Anthropic

Research Engineer, Pretraining Scaling

Anthropic
Apply
14 days ago

Base Salary

$350k - $850k/yr

Responsibilities

  • Own production pretraining pipeline components covering model operations, performance optimization, observability, and reliability.
  • Debug issues across hardware, networking, training dynamics, and evaluation infrastructure.
  • Design and run experiments to improve training efficiency, step time, uptime, and model performance.
  • Respond to on-call incidents during model launches and coordinate solutions across teams.
  • Build and maintain logging, monitoring dashboards, and evaluation infrastructure.
  • Add capabilities such as long-context support and novel architectures to the training codebase.
  • Collaborate with teams across San Francisco and London and document systems, debugging approaches, and lessons learned.

Requirements

  • Hands-on experience training large language models or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.
  • Interest and ability to work across research and engineering in an approximately 50/50 split.
  • Willingness to participate in production on-call work, extended launch hours, and high-pressure incident response.
  • Ability to debug complex, ambiguous problems across multiple layers of the stack.
  • Clear communication and effective collaboration across teams and time zones.
  • Passion for research engineering and responsible AI scaling.
  • Experience with LLM training, large-scale ML frameworks, production ML systems, observability, or evaluation infrastructure is valued.
  • Open-source contributions to frameworks such as open_lm, llm-foundry, or mesh-transformer-jax are valued.
  • Published research in model training, scaling laws, or ML systems is valued.
  • Background as a systems engineer, quant, or in another role combining technical depth and operational excellence is valued.
  • Bachelor’s degree or an equivalent combination of education, training, and/or experience in a relevant field.

Benefits

  • Competitive compensation and benefits, with optional equity donation matching.
  • Generous vacation and parental leave.
  • Flexible working hours and a collaborative office space.
  • This role requires working in-office five days per week in San Francisco.
  • Visa sponsorship may be available, with immigration-lawyer support for sponsored candidates.

Tech Stack

Categories

Anthropic

About Anthropic

501-1,000 employees

We're an AI research company that builds reliable, interpretable, and steerable AI systems. Our first product is Claude, an AI assistant for tasks at any scale. Our research interests span multiple areas including natural language, human feedback, scaling laws, reinforcement learning, code generation, and interpretability.

Contact me