
Site Reliability Engineer
Yes Energy3 months ago
Bucharest, RomaniaSenior
Responsibilities
- Lead incident response from detection through mitigation and recovery, including coordinating responders, restoring service, and driving root-cause remediation.
- Build and improve monitoring, alerting, dashboards, SLOs, runbooks, and escalation processes.
- Operate and troubleshoot Linux and Windows systems across AWS, Azure, OCI, and hybrid or multi-cloud environments.
- Support production web applications, containers, and Kubernetes workloads for reliability, scalability, and availability.
- Diagnose issues involving load balancers, proxies, DNS, networking, firewalls, security groups, and traffic routing.
- Fix Jenkins jobs, CI/CD pipelines, deployment failures, environment issues, and release blockers.
- Partner with Engineering, Security, DBA, and Product Technology Services teams to improve operational readiness and reliability practices.
- Mentor SRE and Systems team members and help establish standards for the growing site reliability function.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
- At least five years of experience supporting mission-critical production infrastructure, SaaS platforms, web applications, or service-oriented systems.
- Deep hands-on AWS experience covering production operations for compute, networking, IAM, storage, load balancing, monitoring, and troubleshooting.
- Proven incident management experience, including high-severity incident leadership, responder coordination, postmortems, root-cause analysis, and corrective actions.
- Experience with containers, Kubernetes, monitoring and alerting systems, CI/CD tooling such as Jenkins and Bitbucket, and operational automation or scripting.
- Working knowledge of Python, PowerShell, Bash, Terraform, CloudFormation, Azure CLI, or AWS CLI.
- Strong Linux and Windows systems administration and production troubleshooting experience.
- Strong communication, collaboration, technical leadership, delegation, mentoring, problem-solving, systems thinking, and prioritization skills.
- Ability to design and maintain scalable infrastructure while applying reliability, access-control, encryption, and data-compliance standards.
Benefits
- Net salary of 14.000–18.000 RON per month.
- Hybrid full-time work arrangement in Bucharest, Romania, with two to three days in the office.
- Private medical insurance, wellness and gym benefits, flexible vacation, and flexible work schedules.
- Company-funded formal and informal professional development.
- Real bonuses and professional growth opportunities.
Categories
Site Reliability