18 days ago
Zürich, SwitzerlandMid Level
Responsibilities
- Own the availability, latency, and performance of production search, retrieval, ingestion, and data-store environments.
- Manage cloud infrastructure across public and private clouds, including networking, identity and access, compute, storage, capacity, quotas, and cost.
- Improve deployment and infrastructure through GitOps-based configuration, infrastructure-as-code, and safe, recoverable production changes.
- Define and instrument SLOs and error budgets and build dashboards, metrics, and alerting for actionable observability.
- Reduce operational toil through automation and internal tooling written in Python and/or Go.
- Provision and standardize new production environments to a defined reliability bar.
- Investigate live incidents, harden systems against recurring failures, and feed reliability learnings back into product engineering and the platform.
Requirements
- At least 3 years of experience in SRE, production or platform engineering, or infrastructure-heavy software engineering.
- Production experience operating container-orchestrated systems and comfort with at least one major public cloud.
- Solid experience with cloud infrastructure, capacity, quotas, and cost management.
- Hands-on observability experience with metrics, logs, and traces, including root-cause analysis of live incidents.
- Fluency with infrastructure-as-code and GitOps.
- Ability to code in Python and/or Go for automation and tooling and to read and patch service code.
- Sound production judgment and the ability to document work clearly for other engineers.
- Preferred experience includes large-scale search or distributed data systems, database operations, dedicated or enterprise-deployed software, or AI and enterprise-search infrastructure in regulated domains.
Benefits
- Ownership of the reliability of systems used by leading law firms at significant scale.
- Hands-on work with distributed search, container orchestration across multiple clouds, infrastructure-as-code, and GitOps.
- Opportunity to establish reliability practices, SLOs, incident processes, and automation from a strong foundation.
- Collaborative engineering team based in Zurich.
- Equal opportunity employer committed to an inclusive, high-performance culture.
