
Site Reliability Engineer, RL Infra
Thinking Machines Lab4 hours ago
Base Salary
$350k - $475k/yr
Responsibilities
- Own end-to-end reliability for RL infrastructure, including observability, SLOs, incident response, and continuous improvement.
- Design monitoring across rollout generation, environment and tool execution, reward computation, and trainer-sampler weight synchronization.
- Drive incident response, recovery, postmortems, and systematic prevention of recurring RL infrastructure failures.
- Harden multi-tenant isolation and resource scheduling for concurrent LoRA-based workloads on shared GPU clusters.
- Reduce the impact of long-tailed rollouts, stale or off-policy data, and stuck trajectories on training stability.
- Build checkpointing and recovery systems for long-running RL jobs.
- Collaborate with security teams to address production vulnerabilities across the RL infrastructure stack.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
- Experience with distributed systems, cloud infrastructure, or site reliability engineering.
- Proficiency writing software, tooling, and automation to solve reliability problems.
- Experience with production incident response, postmortems, and systematic reliability improvement.
- Strong communication and cross-functional coordination experience with engineering and research teams.
- Preferred: deep experience operating production cloud services at scale.
- Preferred: experience with distributed training frameworks and RL infrastructure such as rollout or inference serving, reward pipelines, and asynchronous training systems.
- Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
- Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters for heterogeneous GPU workloads.
Benefits
- Based in San Francisco, California.
- Visa sponsorship is available, with support through the visa process for qualified candidates.
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Tech Stack
Categories
Site Reliability