Anthropic

Staff Software Engineer, Node Infra

Anthropic
Apply
5 months ago
London, United KingdomStaff+

Responsibilities

  • Own the technical strategy and roadmap for node lifecycle management, including ingestion, bring-up, health checking, and automated repair.
  • Drive cross-team initiatives to build and scale AI clusters across multiple clouds and accelerator families.
  • Design and operate systems that detect, isolate, and remediate unhealthy hardware automatically.
  • Define infrastructure architecture and ensure the hardest technical problems are solved directly or through others.
  • Work with cloud providers and internal research, inference, and product teams on compute, data, and infrastructure strategy.
  • Establish and evolve incident response, postmortem, and on-call practices.
  • Mentor and coach engineers and support their professional growth.

Requirements

  • Deep expertise in distributed systems, reliability, and cloud platforms, including technologies such as Kubernetes, infrastructure as code, AWS, GCP, or Azure.
  • Strong proficiency in at least one systems language such as Rust, Go, or Python, plus Terraform infrastructure-as-code proficiency.
  • Hands-on experience with GPUs, TPUs, or Trainium machine learning accelerators.
  • Track record leading complex, multi-quarter technical initiatives across multiple teams or systems.
  • Ability to build alignment across senior stakeholders and communicate effectively at all levels.
  • Preferred: 8+ years of software engineering experience, including technical leadership and direction-setting for a team.
  • Preferred: Experience managing hyperscale compute infrastructure with 10,000+ nodes, including capacity management and efficiency.
  • Preferred: Expertise in Kubernetes internals, cluster orchestration systems, or node provisioning pipelines.
  • Preferred: Low-level systems experience with kernels, virtualization, device drivers, firmware, or hardware health and diagnostics daemons.
  • Preferred: Familiarity with EFA, RDMA, or InfiniBand for distributed machine learning workloads.
  • Preferred: Demonstrated production reliability ownership for high-throughput, latency-sensitive systems.
  • Preferred: Contributions to open-source projects such as Kubernetes, the Linux kernel, or container runtimes.
  • Preferred: Ability to understand systems design tradeoffs and rapidly evolving software systems.
  • Bachelor’s degree in a relevant field or an equivalent combination of education, training, and experience.

Benefits

  • Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more.
  • Visa sponsorship may be available, with immigration lawyer support.
  • Competitive compensation and benefits, excluding the stated salary from this list.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Office space for collaboration.
Anthropic

About Anthropic

5,001-10,000 employees

Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.

Contact me