Anthropic

Research Engineer / Research Scientist, RL Frontiers

Anthropic
Apply
3 hours ago
Seattle, WA, USA +2 moreSenior
H1B sponsor

Base Salary

$500k - $850k/yr

Responsibilities

  • Study how reinforcement learning training and sampling scale with model size, context length, and compute.
  • Develop model architectures and RL algorithms and make them efficient at frontier scale.
  • Scale promising small-scale results to frontier runs and diagnose numerical, algorithmic, and systemic differences.
  • Build fast, reproducible experimental infrastructure for comparing architecture and algorithm variants.
  • Own end-to-end performance of large RL runs from research code through hardware.
  • Build performance and cost models for proposed architecture and algorithm changes.
  • Investigate training instabilities, divergence, and throughput regressions and identify root causes.

Requirements

  • Deep familiarity with modern transformer language models, including architecture, training dynamics, and large-scale optimization.
  • Hands-on experience training large models in distributed settings, including data, tensor, and pipeline parallelism tradeoffs.
  • A record of original technical work in ML training or systems through research, open source, or production impact.
  • Ability to design rigorous large-scale experiments with baselines, ablations, and statistical rigor.
  • Ability to reason quantitatively about model or algorithm compute, memory, and communication costs.
  • Strong programming skills in Python and JAX or PyTorch, with the ability to modify code across the stack.
  • Preferred qualifications include research in reinforcement learning, optimization, or large-scale training; RL algorithms for language models; scaling laws; transformer architecture modification; large accelerator fleets; training numerics; GPU or TPU performance; and C++ or Rust.

Benefits

  • Annual salary range of $500,000—$850,000 USD.
  • Visa sponsorship is offered with an effort caveat: Anthropic cannot successfully sponsor every role and candidate but will make every reasonable effort after making an offer.
  • Hybrid policy currently expects staff to work from an Anthropic office at least 25% of the time, with some roles requiring more office time.
  • Competitive benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office collaboration space.

Categories

AI ResearchML Engineering
Anthropic

About Anthropic

5,001-10,000 employees

Anthropic builds large language models and the Claude AI assistant for developers and enterprises, offered via API access and enterprise plans. Founded in 2021 and headquartered in San Francisco, it distributes Claude through its own platform and via partners such as Amazon Bedrock and Google Cloud’s Vertex AI. Its work emphasizes model reliability, interpretability, and practical tooling for tasks like coding assistance, analysis, and customer support automation.

Contact me