
Site Reliability Engineer (SRE)
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Own end-to-end reliability across CI/CD flows, production observability, and incident response.
- Define service-level objectives for distributed training systems, balancing job completion reliability, scheduling latency, and development velocity.
- Design and implement monitoring and observability across the full training path.
- Lead incident response, recovery, incident reviews, and systematic prevention of recurring issues.
- Harden multi-tenant isolation and resource scheduling for LoRA-based workload co-scheduling without compromising reliability or data separation.
- Collaborate with security teams to address production vulnerabilities.
Requirements
- Bachelor's degree or equivalent experience in computer science, engineering, or a similar field.
- Experience with distributed systems, cloud infrastructure, or site reliability engineering.
- Proficiency writing software, tooling, and automation to solve reliability problems.
- Experience with production incident response, postmortems, and systematic reliability improvement.
- Strong communication and cross-functional coordination skills across engineering and research teams.
- Preferred: deep experience operating production cloud services at scale.
- Preferred: experience with distributed training frameworks and infrastructure failure modes in training.
- Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
- Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters handling heterogeneous GPU workloads.
Benefits
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.