14 days ago
Base Salary
$320k - $485k/yr
Responsibilities
- Design, build, and deploy backend services supporting safety-critical token sampling and generation.
- Own and operate production serving infrastructure across 1P, AWS Bedrock, GCP Vertex, and other deployment platforms.
- Define and maintain SLOs and build observability, alerting, automated validation, and incident-response systems.
- Participate in on-call and operational-duty rotations for service incidents, model provisioning, and time-sensitive launches.
- Build automation, tooling, and self-service workflows that reduce operational toil and manual deployment work.
- Maintain a full-provenance safety registry tracking production deployments, models, timing, and ownership.
- Partner with ML researchers to productionize safety techniques into reliable, scalable systems.
- Contribute to platform-agnostic deployment tooling that brings third-party platforms to operational parity with first-party systems.
Requirements
- Proficiency in Python is required; Rust experience is a plus.
- Experience designing, building, and operating high-QPS systems at global scale.
- Strong knowledge of distributed systems, including replication, consistency tradeoffs, failure modes, and SLO management under load.
- Meaningful production on-call and incident-response experience, including postmortem-driven improvements.
- Hands-on experience deploying and operating AWS and GCP at scale.
- Experience building infrastructure as a platform with abstractions and systems used by other engineers.
- Strong candidates may have 8+ years of industry software engineering experience.
- Experience with deployment and rollout systems, canary analysis, automated validation, or progressive rollout controls is valued.
- A demonstrated history of reducing operational toil through automation and moving teams from manual deployments to self-service pipelines is valued.
- Familiarity with LLM inference systems and transformer-based model operations is valued.
Benefits
- Hybrid work with staff expected to work from an office at least 25% of the time, with some roles requiring more office time.
- Visa sponsorship is available for eligible roles and candidates.
- Competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours, and collaborative office space.
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.
