
AI Infrastructure Engineer
Thinking Machines Lab4 hours ago
San Francisco, CA, USA or New York, NY, USAMid Level
H1B Sponsor
Base Salary
$350k - $475k/yr
Responsibilities
- Own the reliability, performance, and uptime of large-scale post-training and reinforcement learning jobs.
- Partner with research teams during active model runs to debug failures and accelerate iteration.
- Debug distributed-system failures across accelerators, networking, storage, schedulers, and training frameworks.
- Build monitoring, alerting, automated recovery, checkpointing, fault tolerance, and scheduling improvements.
- Develop internal tools that reduce operational toil and improve cluster utilization.
- Participate in an on-call rotation and write postmortems that lead to permanent infrastructure fixes.
Requirements
- At least 4 years of experience as a production, site reliability, or infrastructure engineer operating large-scale distributed systems in production.
- Experience debugging complex networking, hardware, kernel, or scheduler failures.
- Strong software engineering skills in Python and/or Go/C++.
- Solid Linux systems-internals and networking knowledge.
- Experience owning production systems and participating in on-call rotations.
- Preferred experience operating GPU or TPU training clusters at scale.
- Preferred familiarity with RLHF, PPO, DPO, reward model serving, rollout generation, and mixed training/inference workloads.
- Preferred experience with PyTorch, Ray, Slurm, Kubernetes, InfiniBand, RDMA, NCCL, and ML-training observability tooling.
Benefits
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- Participation in an on-call rotation is required.
Categories
ML EngineeringSite Reliability