
Site Reliability Engineer, Model Post Training
Thinking Machines Lab4 hours ago
Base Salary
$350k - $475k/yr
Responsibilities
- Own end-to-end reliability for model post-training workloads across CI/CD, observability, and incident response.
- Define SLOs for distributed training systems while balancing completion reliability, scheduling latency, and development velocity.
- Design monitoring and observability across data loading, training passes, optimization, checkpointing, and sampling.
- Lead incident response, incident reviews, and systematic reliability improvements for cross-team platform issues.
- Harden multi-tenant isolation and resource scheduling for LoRA-based workloads.
- Set reliability direction for newly adopted post-training methods in partnership with research teams.
- Build checkpointing and recovery systems for long-running training jobs.
- Mentor engineers on production ML reliability practices and collaborate with security teams on vulnerabilities.
Requirements
- At least 7 years of experience in distributed systems, cloud infrastructure, or site reliability engineering, including direct on-call ownership of a live production service.
- Hands-on experience operating production machine learning systems for training, fine-tuning, or inference.
- Proficiency writing reliability software and automation, typically in Python and/or Go or Rust.
- Experience leading production incident response, postmortems, and systematic reliability improvements.
- Strong communication skills and experience setting technical direction across engineering and research teams.
- Preferred: deep experience operating cloud services at scale.
- Preferred: knowledge of SFT, RL including PPO or GRPO, DPO, or knowledge distillation.
- Preferred: experience with distributed training frameworks such as PyTorch FSDP/DDP, Megatron, or DeepSpeed.
- Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
- Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters for heterogeneous GPU workloads.
- Preferred: experience establishing reliability standards or SRE practices for an ML platform from the ground up.
Benefits
- Generous health, dental, and vision benefits.
- Unlimited PTO and paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
Tech Stack
Categories
Site Reliability