Anthropic

Research Engineer, Machine Learning (RL Velocity)

Anthropic
Apply
5 months ago
London, United KingdomSenior

Responsibilities

  • Build and improve RL training infrastructure used by researchers daily.
  • Identify and remove bottlenecks through debugging, profiling, and rearchitecting the RL stack.
  • Partner with researchers and adjacent engineering teams to understand pain points and ship productivity-enhancing tooling.
  • Own the reliability and performance of research runs end to end.
  • Contribute to design decisions shaping RL at scale.

Requirements

  • Strong software engineering fundamentals and a track record of building performant, reliable systems.
  • Experience with ML infrastructure, distributed systems, or research tooling.
  • Comfort operating across the stack, including low-level performance work and RL algorithms.
  • Experience with large-scale distributed training in RL, pre-training, or post-training is advantageous.
  • Familiarity with JAX, PyTorch, or similar ML frameworks is advantageous.
  • A bachelor’s degree or equivalent combination of education, training, and/or experience in a relevant field is required.
  • Years of experience will correlate with the internal job-level requirements.

Benefits

  • Annual compensation is £370,000–£630,000 GBP.
  • Hybrid policy requires staff to work from an office at least 25% of the time, with some roles requiring more.
  • Visa sponsorship is available, with immigration-lawyer support.
  • Competitive benefits include optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.
  • Applications are reviewed on a rolling basis with no stated deadline.

Tech Stack

Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.

Contact me