7 months ago
Responsibilities
- Develop SLOs for large language model serving and training systems.
- Design and implement monitoring for availability, latency, and other important metrics.
- Design and implement highly available language-model serving infrastructure for millions of external customers and high-traffic internal workloads.
- Develop and manage automated failover and recovery across multiple regions and cloud providers.
- Lead critical AI-service incident response and drive systematic improvements after incidents.
- Build and maintain cost-optimization systems focused on GPU, TPU, and Trainium utilization and efficiency.
Requirements
- Extensive experience with distributed-systems observability and monitoring at scale.
- Understanding of AI infrastructure, including model serving, batch inference, and training pipelines.
- Proven experience implementing and maintaining SLO/SLA frameworks for business-critical services.
- Ability to work with traditional service metrics and AI-specific metrics such as model performance and training convergence.
- Experience with chaos engineering and systematic resilience testing.
- Ability to bridge ML engineering and infrastructure teams, with strong communication skills.
- Experience operating large-scale model training or serving infrastructure, potentially at more than 1,000 GPUs, is advantageous.
- Experience with GPUs, TPUs, or Trainium, ML-specific networking such as RDMA and InfiniBand, AI observability tools, model deployment strategies, or open-source infrastructure and ML tooling is advantageous.
- Bachelor’s degree in a related field or equivalent experience.
Benefits
- Annual compensation range of £255,000–£325,000 GBP.
- Hybrid policy requiring staff to work from an office at least 25% of the time.
- Visa sponsorship may be available, with immigration-lawyer support.
- Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Categories
DevOpsSite Reliability
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.