7 months ago
Responsibilities
- Develop SLOs for large language model serving and training systems while balancing availability, latency, and development velocity
- Design and implement monitoring systems for availability, latency, and other important metrics
- Help design and implement highly available language model serving infrastructure for millions of external customers and high-traffic internal workloads
- Develop and manage automated failover and recovery systems across multiple regions and cloud providers
- Lead critical AI service incident response and drive systematic improvements after incidents
- Build and maintain cost-optimization systems focused on GPU, TPU, and Trainium utilization and efficiency
Requirements
- Extensive experience with distributed systems observability and monitoring at scale
- Understanding of the operational challenges of AI infrastructure, including model serving, batch inference, and training pipelines
- Proven experience implementing and maintaining SLO/SLA frameworks for business-critical services
- Ability to work with traditional metrics such as latency and availability as well as AI-specific metrics such as model performance and training convergence
- Experience with chaos engineering and systematic resilience testing
- Ability to bridge ML engineering and infrastructure teams
- Excellent communication skills
- Bachelor’s degree in a related field or equivalent experience
- Experience operating large-scale model training or serving infrastructure exceeding 1,000 GPUs is a strong-candidate qualification
- Experience with GPUs, TPUs, or Trainium is a strong-candidate qualification
- Understanding of ML networking optimizations such as RDMA and InfiniBand is a strong-candidate qualification
- Expertise in AI-specific observability tools and frameworks is a strong-candidate qualification
- Understanding of ML model deployment strategies and their reliability implications is a strong-candidate qualification
- Contributions to open-source infrastructure or ML tooling are a strong-candidate qualification
Benefits
- Competitive compensation and benefits
- Optional equity donation matching
- Generous vacation and parental leave
- Flexible working hours
- Lovely office space for collaboration
- Hybrid work arrangement requiring staff to be in an office at least 25% of the time
- Visa sponsorship may be available
- Public benefit corporation environment focused on collaborative AI research
Categories
DevOpsSite Reliability
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.