Anthropic

Staff+ Site Reliability Engineer, Safeguards ML Infra

Anthropic
Apply
3 hours ago
Remote, United States +3 moreStaff+
H1B Sponsor

Base Salary

$320k - $485k/yr

Responsibilities

  • Serve as launch captain and safeguards point of contact for new model releases.
  • Deploy and canary new safety classifiers, run post-deployment validations, investigate discrepancies, and exercise rollback authority when necessary.
  • Verify safeguards across 1P, AWS Bedrock, GCP Vertex, and other deployment platforms while detecting and eliminating configuration drift.
  • Convert launch runbooks and manual checks into tooling, continuous validation, and repeatable self-service deployment pipelines.
  • Build and maintain a safeguards registry recording what is running, on which model and platform, and when and by whom it was deployed.
  • Participate in on-call and operational-duty rotations covering incidents, model provisioning, and time-sensitive launches.

Requirements

  • Experience owning production change management at scale, including deployment pipelines, configuration management, or canary analysis.
  • Experience leading high-stakes releases as a launch captain, incident commander, or release owner.
  • Meaningful production on-call and incident-response experience, including postmortem-driven process and tooling improvements.
  • Hands-on experience deploying and operating on AWS and GCP at scale.
  • Proficiency in Python.
  • Strong candidates may have 8+ years of software engineering or site reliability engineering experience.
  • Strong candidates may have experience reducing operational toil, moving teams from manual deployments to self-service pipelines, and running production-readiness reviews across multiple teams.
  • Rust experience and familiarity with LLM inference systems and transformer-based model operations are preferred.
  • Bachelor’s degree or an equivalent combination of education, training, and/or experience in a relevant field.

Benefits

  • Annual salary range of $320,000–$485,000 USD.
  • Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more office attendance.
  • Visa sponsorship may be available, with immigration-lawyer support.
  • Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.

Categories

DevOpsSite Reliability
Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.

Contact me