Thinking Machines Lab

Site Reliability Engineer, Model Post Training

Thinking Machines Lab
Apply
4 hours ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Own end-to-end reliability for model post-training workloads across CI/CD, observability, and incident response.
  • Define SLOs for distributed training systems while balancing completion reliability, scheduling latency, and development velocity.
  • Design monitoring and observability across data loading, training passes, optimization, checkpointing, and sampling.
  • Lead incident response, incident reviews, and systematic reliability improvements for cross-team platform issues.
  • Harden multi-tenant isolation and resource scheduling for LoRA-based workloads.
  • Set reliability direction for newly adopted post-training methods in partnership with research teams.
  • Build checkpointing and recovery systems for long-running training jobs.
  • Mentor engineers on production ML reliability practices and collaborate with security teams on vulnerabilities.

Requirements

  • At least 7 years of experience in distributed systems, cloud infrastructure, or site reliability engineering, including direct on-call ownership of a live production service.
  • Hands-on experience operating production machine learning systems for training, fine-tuning, or inference.
  • Proficiency writing reliability software and automation, typically in Python and/or Go or Rust.
  • Experience leading production incident response, postmortems, and systematic reliability improvements.
  • Strong communication skills and experience setting technical direction across engineering and research teams.
  • Preferred: deep experience operating cloud services at scale.
  • Preferred: knowledge of SFT, RL including PPO or GRPO, DPO, or knowledge distillation.
  • Preferred: experience with distributed training frameworks such as PyTorch FSDP/DDP, Megatron, or DeepSpeed.
  • Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
  • Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters for heterogeneous GPU workloads.
  • Preferred: experience establishing reliability standards or SRE practices for an ML platform from the ground up.

Benefits

  • Generous health, dental, and vision benefits.
  • Unlimited PTO and paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.

Categories

Site Reliability
Thinking Machines Lab

About Thinking Machines Lab

11-50 employees
Contact me