Thinking Machines Lab

RL Systems - Infra

Thinking Machines Lab
Apply
2 hours ago

Base Salary

$425k - $475k/yr

Responsibilities

  • Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines and workloads.
  • Develop monitoring and observability tools supporting uptime, debuggability, and reproducibility.
  • Collaborate with researchers to turn algorithmic ideas into production-grade training pipelines.
  • Build evaluation and benchmarking infrastructure for model helpfulness, safety, and factuality.
  • Share technical learnings through documentation, open-source libraries, or technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
  • Strong engineering skills, including the ability to write performant, maintainable code and debug complex codebases.
  • Understanding of deep learning frameworks and their underlying system architectures, including PyTorch or JAX.
  • Experience with large-scale language models, reinforcement learning workloads, high-performance or reliability engineering, distributed training frameworks, cluster orchestration, monitoring, observability, or large-scale ML research and infrastructure is preferred.
  • Experience with methods such as PPO, DPO, RLHF, or reward modeling is preferred.

Benefits

  • The role is based in San Francisco, California.
  • Annual base salary is $425,000–$475,000 USD.
  • Visa sponsorship is available for qualified candidates, subject to the company’s stated visa-process caveat.
  • Benefits include health, dental, and vision coverage, unlimited PTO, paid parental leave, and relocation support as needed.
  • This is an evergreen role reviewed on an ongoing basis, and applicants are asked not to reapply more than once every six months.

Tech Stack

GrafanaKubernetesPrometheusPyTorch
Thinking Machines Lab

About Thinking Machines Lab

201-500 employees

Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.

Contact me