2 months ago
Zürich, SwitzerlandMid Level
Responsibilities
- Lead technical responses to production incidents, coordinate resolution across engineering teams, and maintain clear communication.
- Investigate root causes and convert operational learnings into improvements that strengthen platform reliability and resilience.
- Build automation, internal tooling, and operational workflows that improve engineering productivity, deployment safety, and operational efficiency.
- Partner with product engineers to improve the reliability, scalability, and performance of new features before production.
- Help establish engineering best practices for reliability, testing, deployments, observability, and production operations.
Requirements
- Have 4+ years of experience as a Software Engineer, Production Engineer, Site Reliability Engineer, Infrastructure Engineer, or similar role operating production systems at scale.
- Be able to lead technical incident response calmly and communicate and collaborate effectively across engineering teams.
- Be comfortable debugging complex problems across application code, Kubernetes, networking, databases, and cloud infrastructure.
- Have hands-on experience with Kubernetes, containers, Linux, cloud platforms, and modern observability tooling.
- Be comfortable writing software in Go, Python, Rust, Java, or TypeScript.
Benefits
- Primarily work on-site at an office five minutes from Zurich central station.
- Work across distributed infrastructure, Kubernetes, AI systems, and customer-facing applications.
- Solve technically challenging production problems and build tools and infrastructure for the engineering organization.
- Receive a competitive compensation and equity package.
- Join a collaborative, fast-paced, inclusive engineering culture.
Categories
DevOpsSite Reliability
