about 2 hours ago
Berlin, GermanySenior
Responsibilities
- Build and maintain infrastructure automation and infrastructure-as-code at scale.
- Identify and lead large-scale cross-cutting reliability initiatives.
- Design, build, and improve infrastructure components for reliability and observability.
- Define and drive SLOs, error budgets, and alerting standards.
- Participate in on-call rotation and improve on-call experience.
- Collaborate with software engineering teams to embed reliability practices.
Requirements
- 5+ years of hands-on experience in a Site Reliability Engineering role.
- Proven experience with cloud platforms such as AWS, GCP, or Azure.
- Strong experience with Kubernetes and its deployment strategies.
- Experience with infrastructure as code, particularly Terraform.
- Implemented and operated SLIs, SLOs, and error budgets in production.
- Experience managing on-call rotations and leading incident response.
- Familiarity with GitOps workflows like ArgoCD and Helm.
- Proficient in at least one scripting or programming language.
- Fluent in English.
Benefits
- Fully paid Deutschlandticket for public transport.
- 28 vacation days plus an additional day for each full calendar year of employment.
- Work from abroad for up to 10 days per year.
- Company health insurance with supplementary benefits.
- Company pension scheme with employer subsidy.
- Doctolib Parent Care program with additional parental leave.
- Enrollment in long-term employee value sharing plan.
- Free mental health and coaching services.
- Subsidized sports membership.
- Flexible workplace policy with hybrid work options.
- Relocation support for international mobility.
- Access to AI tools for coding and development.
