Thinking Machines Lab

AI Infrastructure Engineer

Thinking Machines Lab
Apply
4 hours ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor

Base Salary

$350k - $475k/yr

Responsibilities

  • Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning jobs.
  • Partner with research teams during active model runs to debug failures and accelerate iteration.
  • Debug distributed-system failures across accelerators, networking, storage, schedulers, and training frameworks.
  • Build monitoring, alerting, automated recovery, checkpointing, fault tolerance, and scheduling improvements.
  • Develop internal tools that reduce operational toil and improve cluster utilization.
  • Participate in an on-call rotation and write postmortems that lead to permanent infrastructure fixes.

Requirements

  • At least 4 years of experience as a production, site reliability, or infrastructure engineer operating large-scale distributed systems in production.
  • Experience debugging complex networking, hardware, kernel, or scheduler failures.
  • Strong software engineering skills in Python and/or Go/C++.
  • Solid Linux systems-internals and networking knowledge.
  • Experience owning production systems and participating in on-call rotations.
  • Preferred experience operating GPU or TPU training clusters at scale.
  • Preferred familiarity with RLHF, PPO, DPO, reward model serving, rollout generation, and mixed training/inference workloads.
  • Preferred experience with PyTorch, Ray, Slurm, Kubernetes, InfiniBand, RDMA, NCCL, and ML-training observability tooling.

Benefits

  • Health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.
  • Participation in an on-call rotation is required.

Categories

ML EngineeringSite Reliability
Thinking Machines Lab

About Thinking Machines Lab

11-50 employees
Contact me