5 months ago
London, United KingdomStaff+
Responsibilities
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management, including provisioning, updates, and decommissioning.
- Partner across teams to ingest new compute capacity on time and align on physical build-out and high-bandwidth inter-cluster connectivity.
- Collaborate with security owners to provision clusters securely by default.
- Define and drive strategy for cluster scalability, homogeneity, and fault tolerance.
- Work with cloud providers and internal research, inference, and product teams on long-term compute, data, and infrastructure strategy.
- Establish and evolve incident response, postmortem, and on-call operational practices.
- Support engineer growth through technical mentorship and coaching.
Requirements
- Deep expertise in distributed systems, reliability, and cloud platforms, including Kubernetes and infrastructure as code across AWS, GCP, or Azure.
- Strong proficiency in at least one systems language such as Rust, Go, or Python, plus infrastructure-as-code proficiency with Terraform.
- Track record of leading complex, multi-quarter technical initiatives across multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively at all levels.
- Preferred: experience operating large-scale compute infrastructure at hyperscale, including 100+ clusters and 10K+ nodes.
- Preferred: depth in Kubernetes internals, cluster provisioning and management systems, or cluster orchestration systems such as Mesos or Borg-like systems.
- Preferred: experience with cloud networking, cluster and host networking, cluster security, Terraform and Atlantis, and workflow orchestration with Temporal or Argo Workflows.
- Preferred: 8+ years of software engineering experience, including time as a technical lead setting direction for a team.
- Bachelor’s degree in a relevant field or an equivalent combination of education, training, and/or experience.
Benefits
- Annual compensation is £325,000–£485,000 GBP.
- Hybrid policy requires staff to work from an office at least 25% of the time, with some roles requiring more.
- Visa sponsorship is available where feasible, with immigration-lawyer support.
- Benefits include competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
Tech Stack
Categories
About Anthropic
Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.
