Anthropic

Research Engineer, Machine Learning (Reinforcement Learning)

Anthropic
Apply
7 months ago
London, United KingdomSenior
H1B Sponsor

Responsibilities

  • Architect and optimize reinforcement-learning infrastructure, including training abstractions and distributed experiment management across GPU clusters.
  • Design, implement, and test novel training environments, evaluations, and reinforcement-learning methodologies.
  • Create agentic models using tool use for computer use, autonomous software generation, mathematics, and other open-ended tasks.
  • Improve system performance through profiling, optimization, benchmarking, caching, and distributed-systems debugging.
  • Develop automated testing frameworks, clean APIs, and scalable infrastructure for AI research.
  • Collaborate with researchers, engineers, alignment teams, frontier red teams, and production-training teams to implement research at scale.

Requirements

  • Proficiency in Python and asynchronous or concurrent programming, including frameworks such as Trio.
  • Experience with machine-learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Industry experience in machine-learning research and the ability to balance research exploration with engineering implementation.
  • Strong systems-design, communication, code-quality, testing, and performance skills.
  • Familiarity with LLM architectures, training methodologies, reinforcement-learning techniques and environments, virtualization, or sandboxed code-execution environments is beneficial.
  • Experience with Kubernetes, distributed systems, or high-performance computing is beneficial.
  • Experience with Rust and/or C++ is beneficial.
  • Formal certifications, academic research experience, and publication history are not required.
  • At least a bachelor's degree in a related field or equivalent experience.

Benefits

  • Competitive benefits package with generous vacation and parental leave.
  • Flexible working hours, optional equity donation matching, and collaborative office space.
  • Hybrid policy requiring staff to work from an office at least 25% of the time.
  • Visa sponsorship may be available, with immigration-lawyer support.

Tech Stack

Categories

AI ResearchML Engineering
Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.

Contact me