GrepJob
Thinking Machines Lab

Research Engineer, Infrastructure, RL Systems

Thinking Machines Lab
Apply
5 days ago
San Francisco, CA, USAMid Level
H1B Sponsor

Base Salary

$350k - $475k/yr

Responsibilities

  • Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
  • Improve the reliability, scalability, and throughput of distributed reinforcement learning training pipelines.
  • Develop monitoring and observability tools that support uptime, debuggability, and reproducibility.
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines.
  • Build evaluation and benchmarking infrastructure for model helpfulness, safety, and factuality.
  • Share technical learnings through internal documentation, open-source libraries, or technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
  • Strong engineering skills, including the ability to write performant, maintainable code and debug complex codebases.
  • Understanding of deep learning frameworks and their underlying system architectures, including PyTorch or JAX.
  • Experience training or supporting language models with tens of billions of parameters or more is preferred.
  • Experience with reinforcement learning workloads such as PPO, DPO, RLHF, or reward modeling is preferred.
  • Background in high-performance or reliability engineering, distributed training frameworks, or cluster orchestration such as Kubernetes or Slurm is preferred.
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, or OpenTelemetry is preferred.
  • Contributions to large-scale machine learning research or infrastructure, open-source frameworks, or performance optimization efforts are preferred.

Benefits

  • Health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.
  • This is an evergreen role reviewed on an ongoing basis.

Tech Stack

GrafanaKubernetesPrometheusPyTorch