
Research Engineer, Infrastructure, Numerics
Thinking Machines Lab2 months ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design and optimize distributed training infrastructure for large-scale LLMs across multi-GPU and multi-node environments.
- Implement and evaluate low-precision numerics such as BF16, MXFP8, and NVFP4.
- Develop kernels and communication primitives using hardware support for mixed- and low-precision arithmetic.
- Collaborate with research teams to co-design model architectures and training recipes around emerging numeric formats and stability constraints.
- Prototype and benchmark data, tensor, and pipeline parallelism strategies with precision-adaptive computation and quantized communication.
- Contribute to internal orchestration and monitoring systems for efficient, reproducible distributed experiments.
- Share technical learnings through documentation, open-source libraries, or technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
- Understanding of deep learning frameworks and their underlying system architectures, including PyTorch and JAX.
- Strong engineering skills and the ability to write performant, maintainable code and debug complex codebases involving floating-point numerics, low-precision arithmetic, and distributed systems.
- Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, or Megatron-LM is preferred.
- Experience implementing FP8, INT8, or block-floating-point formats and understanding their numerical trade-offs is preferred.
- Contributions to open-source deep learning infrastructure such as PyTorch, DeepSpeed, or XLA are preferred.
- Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models are preferred.
- Experience training and supporting large-scale AI models is preferred.
- A track record of improving research productivity through infrastructure or process improvements is preferred.
Benefits
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- This is an evergreen role reviewed on an ongoing basis.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.