
Site Reliability Engineer (m/f/d)
Codesphere4 months ago
Remote, Germany or Munich, GermanyMid Level
Responsibilities
- Define and enforce SLOs, SLIs, and SLAs across production.
- Monitor system health, plan capacity, and automate deployments, patching, and infrastructure provisioning.
- Diagnose and resolve production incidents, including participating in 24/7 on-call coverage.
- Lead post-mortems, implement preventive improvements, and maintain runbooks and escalation procedures.
- Manage cloud infrastructure through IaC and design and maintain CI/CD pipelines.
- Drive scalability, fault tolerance, disaster recovery, and security compliance.
- Partner with development teams on production readiness, Shift Left practices, and error budget management.
Requirements
- Proven experience in an SRE, DevOps, or platform engineering role with hands-on production ownership.
- Strong knowledge of Kubernetes, Terraform, and Ansible.
- Familiarity with Ceph or comparable distributed storage systems.
- Experience with SLOs, SLIs, error budgets, and CI/CD pipeline design.
- A degree in a relevant field or a comparable qualification.
- Strong debugging and incident response skills, with the ability to remain calm and structured under pressure.
- Good communication skills and the ability to translate operational concerns into guidance for development teams.
- Go development experience is a plus.
Benefits
- 32 days of paid time off, including Christmas Eve and New Year's Eve.
- Meal allowance of up to 15 digital vouchers per month, totaling over €100 net.
- Hybrid work setup with mobile work options and flexibility around core hours.
- Fast-moving environment with substantial ownership and opportunities to learn.
- Tax-free bicycle leasing through Job-Rad.
- Gym access at the Karlsruhe office.
- Employee events including team offsites and regular get-togethers.
- Company-supported pension scheme.
- Offices with convenient access to tram and metro stops in Karlsruhe and Munich.
Tech Stack
Categories
Site Reliability