Thinking Machines Lab

Site Reliability Engineer, RL Infra

Thinking Machines Lab
Apply
4 hours ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Own end-to-end reliability for RL infrastructure, including observability, SLOs, incident response, and continuous improvement.
  • Design monitoring across rollout generation, environment and tool execution, reward computation, and trainer-sampler weight synchronization.
  • Drive incident response, recovery, postmortems, and systematic prevention of recurring RL infrastructure failures.
  • Harden multi-tenant isolation and resource scheduling for concurrent LoRA-based workloads on shared GPU clusters.
  • Reduce the impact of long-tailed rollouts, stale or off-policy data, and stuck trajectories on training stability.
  • Build checkpointing and recovery systems for long-running RL jobs.
  • Collaborate with security teams to address production vulnerabilities across the RL infrastructure stack.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
  • Experience with distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software, tooling, and automation to solve reliability problems.
  • Experience with production incident response, postmortems, and systematic reliability improvement.
  • Strong communication and cross-functional coordination experience with engineering and research teams.
  • Preferred: deep experience operating production cloud services at scale.
  • Preferred: experience with distributed training frameworks and RL infrastructure such as rollout or inference serving, reward pipelines, and asynchronous training systems.
  • Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
  • Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters for heterogeneous GPU workloads.

Benefits

  • Based in San Francisco, California.
  • Visa sponsorship is available, with support through the visa process for qualified candidates.
  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.

Tech Stack

Categories

Site Reliability
Thinking Machines Lab

About Thinking Machines Lab

11-50 employees
Contact me