
Software Engineer, Infrastructure
Thinking Machines Lab5 hours ago
Base Salary
$300k - $400k/yr
Responsibilities
- Design, build, and operate distributed systems for large-scale model training and inference across thousands of accelerators.
- Build and maintain orchestration, scheduling, storage, and resource-management infrastructure.
- Improve the reliability, performance, and observability of infrastructure used by research and product teams.
- Partner with researchers and platform engineers to translate infrastructure needs into robust, well-abstracted systems.
- Debug and resolve distributed failures involving networking, storage, compute, and scheduling.
- Write and maintain internal Python and Go libraries and APIs.
Requirements
- Demonstrated expertise designing and developing large-scale distributed systems.
- Strong proficiency in Python and Go.
- Experience building, deploying, and operating production infrastructure at scale.
- Strong understanding of distributed systems fundamentals, including consensus, consistency, fault tolerance, and networking.
- Preferred experience with ML infrastructure such as training orchestration, job schedulers, or distributed storage and data systems.
- Preferred experience operating large-scale GPU or TPU clusters.
- Preferred experience with Kubernetes and infrastructure-as-code.
- Open-source infrastructure contributions are preferred.
- Ability to work autonomously in a fast-changing, early-stage environment.
Benefits
- Based in San Francisco, California.
- Annual salary range of $300,000–$400,000 USD.
- Visa sponsorship is available, with a commitment to work through the visa process for the right fit, though success is not guaranteed for every candidate or role.
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.