3 months ago
Responsibilities
- Define and own the platform’s production-grade reliability practice end to end.
- Build service-level objectives, monitoring, alerting, and observability.
- Own on-call and incident response, including runbooks, escalation, blameless postmortems, and follow-through.
- Lead capacity planning and the operational management of heterogeneous compute backends.
- Set technical direction in collaboration with the platform team and work with hardware and orchestration teams to expose backends reliably.
Requirements
- Strong SRE or production-engineering background running customer-facing systems at scale.
- Fluency with modern operational tooling, including observability stacks, container orchestration, infrastructure-as-code, and CI/CD.
- Experience owning incident response and driving reliability improvements.
- Compliance execution experience.
- Open-source contributions to relevant infrastructure or detailed experience with production systems of significant scale and complexity is valued.
- Early-stage and AI-native company experience is valued.
Benefits
- Competitive salary determined by skills and experience.
- Equity and ownership.
- Private healthcare.
- Visa sponsorship and relocation benefits.
- In-person work at the London office with provided tools, workspace, and setup.
Categories
Site Reliability
