about 2 hours ago
Paris, FranceSenior
Responsibilities
- Build and maintain infrastructure automation and infrastructure-as-code at scale.
- Identify and lead large-scale cross-cutting reliability initiatives.
- Design, build, and improve infrastructure components for reliability and observability.
- Define and drive SLOs, error budgets, and alerting standards.
- Participate in the on-call rotation and improve the on-call experience.
- Collaborate with software engineering teams to embed reliability practices.
Requirements
- 5+ years of hands-on experience in a Site Reliability Engineering role.
- Proven experience with cloud platforms such as AWS, GCP, or Azure.
- Strong experience with Kubernetes and its deployment strategies.
- Experience with infrastructure as code, particularly Terraform.
- Implemented and operated SLIs, SLOs, and error budgets in production.
- Experience managing on-call rotations and leading incident response.
- Familiarity with GitOps workflows like ArgoCD and Helm.
- Proficiency in at least one scripting or programming language.
- Fluency in English.
Benefits
- Free comprehensive health insurance for you and your children.
- 25 days of paid vacation per year, plus up to 14 days of RTT.
- Free mental health and coaching services.
- Work from abroad for up to 10 days per year.
- Lunch vouchers worth 8.50 euros per working day.
- 50% reimbursement of your public transport subscription.
- Additional month of leave on top of legal parental leave.
- Relocation support for international mobility.
- Access to the best AI tools for coding and development.
Tech Stack
AWSDatadogGoGoogle Cloud PlatformJavaKotlinKubernetesPrometheusPythonReact NativeRubyRuby on RailsSwiftTerraformTypeScript
