
Research Engineer, Infrastructure, Training Systems
Thinking Machines Lab2 months ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design, implement, and optimize distributed training systems across thousands of GPUs and nodes.
- Develop high-performance optimizations that maximize training throughput and efficiency.
- Build reusable frameworks and libraries that improve training reproducibility, reliability, and scalability for new model architectures.
- Establish reliability, maintainability, and security standards for rapidly iterated systems.
- Collaborate with researchers and engineers to build scalable infrastructure.
- Share technical learnings through internal documentation, open-source libraries, or technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
- Strong engineering skills with the ability to write performant, maintainable code and debug complex codebases.
- Understanding of deep learning frameworks such as PyTorch and JAX and their underlying system architectures.
- Ability to collaborate effectively with cross-functional partners and subject matter experts.
- Initiative and willingness to work across stacks and teams to deliver projects.
- Preferred experience working on distributed training for very large models to improve stability, reliability, and performance.
- Preferred track record of improving research productivity through infrastructure design or process improvements.
- Preferred contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.
Benefits
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- This is an evergreen role reviewed on an ongoing basis.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.