Wayve

Staff SRE, AI Infrastructure

Wayve
Apply
1 day ago
London, United KingdomStaff+

Responsibilities

  • Own the reliability, availability, and performance of the Model Development Platform and GPU Compute environments.
  • Define and operationalise SLOs, SLIs, error budgets, production-readiness standards, and reliability programs.
  • Improve capacity planning, scaling strategies, resource efficiency, deployment safety, and rollback mechanisms across large GPU-backed clusters.
  • Participate in a 24/7 on-call rotation and lead incident triage, escalation, communications, root cause analysis, and postmortems.
  • Design and operate monitoring, logging, tracing, alerting, and user-centric platform health dashboards.
  • Build automation for cluster operations, training workflows, remediation, scaling, self-healing, CI/CD, and infrastructure guardrails.

Requirements

  • Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
  • Experience operating GPU-backed environments, large-scale ML infrastructure, or complex compute clusters.
  • Experience running model training or inference pipelines in production, including MLOps workloads.
  • Strong Kubernetes experience operating production clusters.
  • Hands-on experience running production workloads in AWS, GCP, or Azure.
  • Strong Linux fundamentals and proficiency in at least one scripting or systems language such as Python, Go, or C++.
  • Deep troubleshooting experience across networking, storage, distributed systems, and performance at scale.
  • Experience designing and operating observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Clear communication skills, including leading incidents, writing postmortems, and influencing reliability improvements.
  • Familiarity with infrastructure-as-code such as Terraform and secure cloud production environments is desirable.
  • Experience defining and running SLOs and SLIs, establishing SRE processes from scratch, or growing a Cloud SRE function is desirable.

Benefits

  • Full-time employment.
  • Hybrid working arrangement based in London with two days per week in the office.
  • Inclusive interview experience with accommodations available upon request.

Categories

DevOpsSite Reliability
Wayve

About Wayve

1,001-5,000 employees
Contact me