Wayve

Senior SRE, AI Infrastructure

Wayve
Apply
1 day ago
London, United KingdomSenior

Responsibilities

  • Own the reliability, availability, and performance of the Model Development Platform and GPU Compute environments.
  • Define and operationalise SLOs, SLIs, error budgets, capacity planning, scaling strategies, and production readiness standards.
  • Participate in a 24/7 on-call rotation and lead incident triage, escalation, communications, root cause analysis, and post-incident improvements.
  • Design and operate monitoring, logging, tracing, alerting, dashboards, and other observability systems.
  • Build automation for cluster operations, training workflows, remediation, scaling, self-healing, infrastructure-as-code, and policy-driven guardrails.
  • Improve deployment safety through change management, validation, rollback mechanisms, and CI/CD improvements.
  • Partner with ML, platform, and software teams to reduce operational burden and improve platform reliability.

Requirements

  • Proven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.
  • Strong experience operating production Kubernetes clusters.
  • Hands-on experience running production workloads in AWS, GCP, or Azure.
  • Experience operating complex distributed systems, preferably including compute-heavy or high-performance workloads.
  • Experience working with large compute clusters; AI/ML training or inference experience is strongly preferred.
  • Strong Linux fundamentals and proficiency in at least one scripting or systems language such as Python, Go, or C++.
  • Deep troubleshooting skills across networking, storage, distributed systems, and performance at scale.
  • Experience designing and operating observability stacks such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Clear communication skills, including leading incidents, writing postmortems, and influencing reliability improvements.
  • Experience operating GPU-backed environments, large-scale ML infrastructure, or production model training and inference pipelines is desirable.
  • Familiarity with Terraform, secure cloud production environments, SLOs, SLIs, and reliability programs is desirable.
  • Experience establishing an SRE function from scratch and interest in taking on future leadership responsibilities are desirable.

Benefits

  • Full-time hybrid role based in London with two days per week in the office and the remainder working from home.
  • Inclusive interview process with accommodations and adjustments available upon request.

Categories

Site Reliability
Wayve

About Wayve

1,001-5,000 employees
Contact me