7 months ago
Base Salary
$325k - $485k/yr
Responsibilities
- Develop service-level objectives for large language model serving systems while balancing availability, latency, and development velocity.
- Design and implement monitoring and observability systems across the token path.
- Help design and implement highly available serving infrastructure across multiple regions and cloud providers.
- Lead incident response for critical AI services, including recovery, incident reviews, and systematic improvements.
- Support the reliability of safeguard model serving in alignment with Anthropic’s safety commitments.
- Partner cross-functionally with teams across Anthropic to improve reliability across critical serving paths.
Requirements
- Strong background in distributed systems, infrastructure, or reliability engineering.
- Ability to investigate unfamiliar systems during incidents and help drive resolution.
- Holistic understanding of how complex systems compose and where reliability seams exist.
- Strong communication, collaboration, relationship-building, and ownership skills.
- Experience as an SRE, Production Engineer, or in a similar reliability-focused role is valuable.
- Experience operating large-scale model serving or training infrastructure involving more than 1,000 GPUs is valuable.
- Experience with GPUs, TPUs, or Trainium is valuable.
- Understanding of ML networking optimizations such as RDMA and InfiniBand is valuable.
- Experience with AI-specific observability tools and frameworks, chaos engineering, systematic resilience testing, or open-source infrastructure and ML tooling is valuable.
- Bachelor’s degree in a related field or equivalent experience.
Benefits
- Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more office time.
- Visa sponsorship may be available, with immigration-lawyer support.
- Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Categories
DevOpsSite Reliability
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.