Anthropic

Staff+ Software Engineer, Inference Runtime

Anthropic
Apply
3 months ago
Seattle, WA, USA +2 moreStaff+
H1B sponsor

Base Salary

$405k - $485k/yr

Responsibilities

  • Set technical direction, architecture, and roadmap for the shared inference runtime
  • Own and evolve the accelerator-agnostic runtime’s interfaces, internal boundaries, and build structure through hands-on Rust and Python development
  • Design abstractions that allow new models and deployment targets to add specialization without increasing core platform cost
  • Drive accelerator utilization, scheduling, and memory management across GPU, TPU, and Trainium
  • Build validation capabilities using partitioned builds, change-scoped testing, canary and shadow mechanisms, and rollback
  • Partner with Infrastructure on compilers, build systems, and toolchains and determine whether to build or adopt dependencies
  • Mentor engineers through design reviews, code reviews, and direct collaboration without owning headcount
  • Represent the team in cross-organizational efforts across serving, scaling, and accelerator teams

Requirements

  • Deep systems engineering or ML infrastructure experience, including performance profiling, latency and throughput optimization, and large-scale systems debugging
  • Significant software engineering experience with high-performance, large-scale distributed systems serving millions of users
  • Deep experience in at least one accelerator ecosystem, such as CUDA/GPU, TPU, or Trainium/AWS Neuron
  • Experience defining and using engineering metrics, including SLOs, escape rates, release times, latency, or throughput
  • Experience driving technical alignment across organizational boundaries and influencing technical direction without formal authority
  • Strong written and verbal communication skills
  • Preferred: 8+ years of software engineering experience and significant time as a technical lead or platform, inference runtime, or ML infrastructure anchor
  • Preferred: Experience with XLA, Triton, NeuronX, or accelerator driver and firmware management at scale
  • Preferred: Experience with shadow traffic, canary populations, automated baseline comparison, and fast rollback
  • Preferred: Experience with deterministic or simulation-based testing for hardware-dependent systems
  • Preferred: Experience with CI/CD systems at scale for accelerator hardware workloads
  • Preferred: Familiarity with Kubernetes-based development and job scheduling environments
  • Preferred: Prior technical lead experience on a developer productivity or platform engineering team at a fast-growing AI/ML company
  • Bachelor’s degree or equivalent combination of education, training, and/or experience in a relevant field

Benefits

  • Hybrid policy requiring staff to be in an office at least 25% of the time
  • Visa sponsorship may be available, with immigration lawyer support
  • Competitive compensation and benefits
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Office space for collaboration
Anthropic

About Anthropic

5,001-10,000 employees

Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.

Contact me