GrepJob
Thinking Machines Lab

Site Reliability Engineer (SRE)

Thinking Machines Lab
Apply
5 days ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Own end-to-end reliability across CI/CD flows, production observability, and incident response.
  • Define service-level objectives for distributed training systems, balancing job completion reliability, scheduling latency, and development velocity.
  • Design and implement monitoring and observability across the full training path.
  • Lead incident response, recovery, incident reviews, and systematic prevention of recurring issues.
  • Harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling without compromising reliability or data separation.
  • Collaborate with security teams to address production vulnerabilities.

Requirements

  • Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
  • Experience with distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software, tooling, and automation to solve reliability problems.
  • Experience with production incident response, postmortems, and systematic reliability improvement.
  • Strong communication and cross-functional coordination skills across engineering and research teams.
  • Preferred: deep experience operating production cloud services at scale.
  • Preferred: experience with distributed training frameworks and infrastructure failure modes in training.
  • Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
  • Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters handling heterogeneous GPU workloads.

Benefits

  • Health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.

Tech Stack

Categories