14 days ago
Base Salary
$320k - $485k/yr
Responsibilities
- Own the technical strategy and roadmap for accelerator node lifecycle management, including ingestion, provisioning, health checking, and automated repair.
- Drive cross-team initiatives to build and scale AI clusters across multiple cloud providers and accelerator families.
- Design and operate systems that detect, isolate, and remediate unhealthy hardware while improving fleet reliability and capacity utilization.
- Define infrastructure architecture and solve complex technical problems directly or through other engineers.
- Partner with cloud providers and internal research, inference, and product teams on compute, data, and infrastructure strategy.
- Establish operational excellence practices including incident response, postmortems, and on-call operations.
- Mentor and coach engineers and support their technical growth.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms such as Kubernetes, AWS, GCP, and Azure.
- Strong proficiency in at least one systems language, such as Rust, Go, or Python, plus proficiency with Terraform for infrastructure as code.
- Hands-on experience with GPUs, TPUs, or Trainium machine learning accelerators.
- Experience leading complex, multi-quarter technical initiatives across multiple teams or systems.
- Ability to align senior stakeholders and communicate effectively at all levels.
- Preferred: 10+ years of software engineering experience, including technical leadership and direction-setting for a team.
- Preferred: experience managing hyperscale compute infrastructure with 10,000+ nodes, including capacity management and efficiency.
- Preferred: depth in Kubernetes internals, cluster orchestration systems, or node provisioning pipelines.
- Preferred: low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
- Preferred: familiarity with EFA, RDMA, or InfiniBand for distributed machine learning workloads.
- Preferred: production reliability experience with high-throughput, latency-sensitive systems and contributions to open-source projects such as Kubernetes, the Linux kernel, or container runtimes.
- Bachelor’s degree in a relevant field, or an equivalent combination of education, training, and/or experience.
Benefits
- Hybrid work arrangement requiring staff to be in an office at least 25% of the time, with some roles requiring more office time.
- Visa sponsorship and immigration-lawyer support may be available.
- Competitive benefits, generous vacation and parental leave, flexible working hours, and office collaboration space.
- Optional equity donation matching.
Categories
About Anthropic
We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.
