3 hours ago
Base Salary
$500k - $850k/yr
Responsibilities
- Design, build, and operate distributed systems for RL training, sampling, and environment execution at scale.
- Optimize scheduling, placement, data movement, storage, networking, coordination, and other system bottlenecks.
- Build fault detection, isolation, recovery, autoscaling, observability, and automated remediation for long-running jobs.
- Diagnose complex failures across hosts and services, remove recurring failure classes, and preserve training correctness.
- Collaborate with researchers and performance engineers, write design documents and incident reports, and create safe operational interfaces for humans and automated tools.
Requirements
- Strong software engineering skills in Python and at least one systems language such as Rust, C++, or Go.
- Experience designing, building, and operating large-scale distributed systems in production.
- Deep understanding of distributed-systems fundamentals, including consistency, coordination, consensus, failure modes, and recovery.
- Ability to analyze throughput, latency, and resource costs across compute, memory, storage, and networking.
- Experience debugging complex, non-localizable failures across many hosts and services.
- Strong written communication skills, including design documents and incident writeups.
- Preferred qualifications include ML training or inference infrastructure, schedulers, autoscalers, resource management, Kubernetes, sandboxed or virtualized execution, high-performance networking, RDMA, collective communication libraries, fleet observability, automated remediation, Trio or asyncio, and familiarity with RL or LLM training workloads.
Benefits
- Annual base salary range of $500,000–$850,000 USD.
- Visa sponsorship is offered with reasonable efforts made after an offer, though sponsorship cannot be successfully provided for every role or candidate.
- Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more office time.
- Competitive benefits including optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
About Anthropic
Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.
