1 day ago
London, United KingdomStaff+
Responsibilities
- Own the reliability, availability, and performance of the Model Development Platform and GPU Compute environments.
- Define and operationalise SLOs, SLIs, error budgets, production-readiness standards, and reliability programs.
- Improve capacity planning, scaling strategies, resource efficiency, deployment safety, and rollback mechanisms across large GPU-backed clusters.
- Participate in a 24/7 on-call rotation and lead incident triage, escalation, communications, root cause analysis, and postmortems.
- Design and operate monitoring, logging, tracing, alerting, and user-centric platform health dashboards.
- Build automation for cluster operations, training workflows, remediation, scaling, self-healing, CI/CD, and infrastructure guardrails.
Requirements
- Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
- Experience operating GPU-backed environments, large-scale ML infrastructure, or complex compute clusters.
- Experience running model training or inference pipelines in production, including MLOps workloads.
- Strong Kubernetes experience operating production clusters.
- Hands-on experience running production workloads in AWS, GCP, or Azure.
- Strong Linux fundamentals and proficiency in at least one scripting or systems language such as Python, Go, or C++.
- Deep troubleshooting experience across networking, storage, distributed systems, and performance at scale.
- Experience designing and operating observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry.
- Clear communication skills, including leading incidents, writing postmortems, and influencing reliability improvements.
- Familiarity with infrastructure-as-code such as Terraform and secure cloud production environments is desirable.
- Experience defining and running SLOs and SLIs, establishing SRE processes from scratch, or growing a Cloud SRE function is desirable.
Benefits
- Full-time employment.
- Hybrid working arrangement based in London with two days per week in the office.
- Inclusive interview experience with accommodations available upon request.
Tech Stack
Categories
DevOpsSite Reliability
