GrepJob
Thinking Machines Lab

Research Engineer, Infrastructure, Training Systems

Thinking Machines Lab
Apply
5 days ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Design, implement, and optimize distributed training systems across thousands of GPUs and nodes for large-scale workloads.
  • Develop high-performance optimizations that maximize training throughput and efficiency.
  • Build reusable frameworks and libraries that improve reproducibility, reliability, and scalability for new model architectures.
  • Establish reliability, maintainability, and security standards for robust systems under rapid iteration.
  • Collaborate with researchers and engineers to build scalable infrastructure.
  • Share learnings through internal documentation, open-source libraries, or technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
  • Strong engineering skills with the ability to write performant, maintainable code and debug complex codebases.
  • Understanding of deep learning frameworks such as PyTorch and JAX and their underlying system architectures.
  • Experience working collaboratively with cross-functional partners and subject matter experts.
  • Preferred: experience with distributed training for the world’s largest models, improving research productivity through infrastructure or process improvements, or contributing to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Benefits

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.
  • This is an evergreen role reviewed on an ongoing basis rather than a guaranteed immediate opening.

Tech Stack

Categories