
Research Engineer, Infrastructure, Training Systems
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design, implement, and optimize distributed training systems across thousands of GPUs and nodes for large-scale workloads.
- Develop high-performance optimizations that maximize training throughput and efficiency.
- Build reusable frameworks and libraries that improve reproducibility, reliability, and scalability for new model architectures.
- Establish reliability, maintainability, and security standards for robust systems under rapid iteration.
- Collaborate with researchers and engineers to build scalable infrastructure.
- Share learnings through internal documentation, open-source libraries, or technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
- Strong engineering skills with the ability to write performant, maintainable code and debug complex codebases.
- Understanding of deep learning frameworks such as PyTorch and JAX and their underlying system architectures.
- Experience working collaboratively with cross-functional partners and subject matter experts.
- Preferred: experience with distributed training for the world’s largest models, improving research productivity through infrastructure or process improvements, or contributing to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.
Benefits
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- This is an evergreen role reviewed on an ongoing basis rather than a guaranteed immediate opening.