
Site Reliability Engineer, Production
Thinking Machines Lab23 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Own end-to-end reliability for Tinker, including production observability, incident response, and reliability improvements.
- Define SLOs for distributed training systems while balancing job completion reliability, scheduling latency, and development velocity.
- Design and implement monitoring and observability across the full training path.
- Lead incident response, incident reviews, and systematic efforts to prevent recurrence.
- Harden multi-tenant isolation and resource scheduling for reliable, efficient LoRA workload co-scheduling.
- Collaborate with security teams to address production vulnerabilities.
- Build software tooling and automation to solve reliability problems.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
- Experience with distributed systems, cloud infrastructure, or site reliability engineering.
- Proficiency writing software, tooling, and automation for reliability problems.
- Experience with production incident response, postmortems, and systematic reliability improvement.
- Strong communication skills and experience coordinating across engineering and research teams.
- Preferred: deep experience operating production cloud services at scale.
- Preferred: experience with distributed training frameworks and infrastructure failure modes in training systems.
- Preferred: experience building checkpoint and recovery systems for long-running distributed jobs.
- Preferred: expertise deploying, operating, debugging, and tuning Kubernetes clusters for heterogeneous GPU workloads.
Benefits
- Based in San Francisco, California.
- Annual salary range of $350,000–$475,000 USD, depending on background, skills, and experience.
- Visa sponsorship is available.
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Tech Stack
Categories
Site Reliability
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.