Anthropic

Research Engineer, Code RL (Reinforcement Learning)

Anthropic
Apply
3 months ago
San Francisco, CA, USA or New York, NY, USASenior
H1B sponsor

Base Salary

$500k - $850k/yr

Responsibilities

  • Design reinforcement-learning environments and coding tasks for software-engineering capabilities.
  • Build reward signals and verifiers that evaluate code quality and correctness.
  • Run training experiments on frontier models and diagnose model improvements or failures.
  • Improve the speed and reliability of research and training pipelines.
  • Advance agentic coding, code correctness, long-horizon autonomous engineering, and high-performance accelerator code.
  • Collaborate with alignment, red-team, and production-training teams to implement research at scale.

Requirements

  • Strong software-engineering skills and deep Python expertise, including asynchronous and concurrent programming.
  • Ability to own systems end to end and debug across the stack.
  • Ability to balance research exploration with engineering implementation and rigorously interpret experimental results.
  • Care for code quality, testing, and performance.
  • Commitment to developing safe and beneficial AI systems.
  • Preferred experience with reinforcement learning, RLHF, post-training, or LLM fine-tuning.
  • Preferred experience building coding agents, code-execution sandboxes, evaluation harnesses, verifiers, or developer tooling.
  • Preferred background in program analysis, testing, verification, compilers, or formal methods.
  • Preferred experience with PyTorch, large-scale distributed training, performance profiling, and ML-systems optimization.
  • Preferred CUDA, GPU, TPU-kernel, virtualization, or sandboxed-code-execution experience.

Benefits

  • Hybrid policy requiring staff to work from an office at least 25% of the time, with some roles requiring more.
  • Visa sponsorship may be available, with immigration-lawyer support.
  • Optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours.
  • Collaborative office space.

Tech Stack

Categories

Anthropic

About Anthropic

5,001-10,000 employees

Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.

Contact me