Anthropic

Senior Software Engineer, AI Reliability Engineering

Anthropic
Apply
7 months ago
Dublin, IrelandSenior

Responsibilities

  • Develop SLOs for large language model serving and training systems while balancing availability, latency, and development velocity
  • Design and implement monitoring systems for availability, latency, and other important metrics
  • Help design and implement highly available language model serving infrastructure for millions of external customers and high-traffic internal workloads
  • Develop and manage automated failover and recovery systems across multiple regions and cloud providers
  • Lead critical AI service incident response and drive systematic improvements after incidents
  • Build and maintain cost-optimization systems focused on GPU, TPU, and Trainium utilization and efficiency

Requirements

  • Extensive experience with distributed systems observability and monitoring at scale
  • Understanding of the operational challenges of AI infrastructure, including model serving, batch inference, and training pipelines
  • Proven experience implementing and maintaining SLO/SLA frameworks for business-critical services
  • Ability to work with traditional metrics such as latency and availability as well as AI-specific metrics such as model performance and training convergence
  • Experience with chaos engineering and systematic resilience testing
  • Ability to bridge ML engineering and infrastructure teams
  • Excellent communication skills
  • Bachelor’s degree in a related field or equivalent experience
  • Experience operating large-scale model training or serving infrastructure exceeding 1,000 GPUs is a strong-candidate qualification
  • Experience with GPUs, TPUs, or Trainium is a strong-candidate qualification
  • Understanding of ML networking optimizations such as RDMA and InfiniBand is a strong-candidate qualification
  • Expertise in AI-specific observability tools and frameworks is a strong-candidate qualification
  • Understanding of ML model deployment strategies and their reliability implications is a strong-candidate qualification
  • Contributions to open-source infrastructure or ML tooling are a strong-candidate qualification

Benefits

  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Lovely office space for collaboration
  • Hybrid work arrangement requiring staff to be in an office at least 25% of the time
  • Visa sponsorship may be available
  • Public benefit corporation environment focused on collaborative AI research

Categories

DevOpsSite Reliability
Anthropic

About Anthropic

5,001-10,000 employees

Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.

Contact me