1 day ago
London, United KingdomSenior
Responsibilities
- Own the reliability, availability, and performance of the Model Development Platform and GPU Compute environments.
- Define and operationalise SLOs, SLIs, error budgets, capacity planning, scaling strategies, and production readiness standards.
- Participate in a 24/7 on-call rotation and lead incident triage, escalation, communications, root cause analysis, and post-incident improvements.
- Design and operate monitoring, logging, tracing, alerting, dashboards, and other observability systems.
- Build automation for cluster operations, training workflows, remediation, scaling, self-healing, infrastructure-as-code, and policy-driven guardrails.
- Improve deployment safety through change management, validation, rollback mechanisms, and CI/CD improvements.
- Partner with ML, platform, and software teams to reduce operational burden and improve platform reliability.
Requirements
- Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
- Strong experience operating production Kubernetes clusters.
- Hands-on experience running production workloads in AWS, GCP, or Azure.
- Experience operating complex distributed systems, preferably including compute-heavy or high-performance workloads.
- Experience working with large compute clusters; AI/ML training or inference experience is strongly preferred.
- Strong Linux fundamentals and proficiency in at least one scripting or systems language such as Python, Go, or C++.
- Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale.
- Experience designing and operating observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry.
- Clear communication skills, including leading incidents, writing postmortems, and influencing reliability improvements.
- Experience operating GPU-backed environments, large-scale ML infrastructure, or production model training and inference pipelines is desirable.
- Familiarity with Terraform, secure cloud production environments, SLOs, SLIs, and reliability programs is desirable.
- Experience establishing an SRE function from scratch and interest in taking on future leadership responsibilities are desirable.
Benefits
- Full-time hybrid role based in London with two days per week in the office and the remainder working from home.
- Inclusive interview process with accommodations and adjustments available upon request.
Tech Stack
Categories
Site Reliability
