GrepJob
Thinking Machines Lab

Research Engineer, Infrastructure, Numerics

Thinking Machines Lab
Apply
5 days ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Design and optimize distributed training infrastructure for large-scale LLMs across multi-GPU and multi-node environments.
  • Implement and evaluate low-precision numerics such as BF16, MXFP8, and NVFP4.
  • Develop kernels and communication primitives using hardware support for mixed- and low-precision arithmetic.
  • Collaborate with research teams to co-design model architectures and training recipes for emerging numeric formats and stability constraints.
  • Prototype and benchmark data, tensor, and pipeline parallelism with precision-adaptive computation and quantized communication.
  • Help design internal orchestration and monitoring systems for efficient, reproducible distributed experiments.
  • Share findings through internal documentation, open-source libraries, or technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
  • Understanding of deep learning frameworks and their underlying system architectures.
  • Strong engineering skills with the ability to write performant, maintainable code and debug complex codebases involving floating-point numerics, low-precision arithmetic, and distributed systems.
  • Familiarity with PyTorch/XLA, DeepSpeed, or Megatron-LM is preferred.
  • Experience implementing FP8, INT8, or block-floating-point formats and understanding their numerical trade-offs is preferred.
  • Prior contributions to open-source deep learning infrastructure, such as PyTorch, DeepSpeed, or XLA, are preferred.
  • Publications, patents, or projects related to numerical optimization, communication-efficient training, or large-model systems are preferred.
  • Experience training and supporting large-scale AI models is preferred.
  • A track record of improving research productivity through infrastructure or process improvements is preferred.
  • The role requires effective collaboration across cross-functional partners and subject matter experts.

Benefits

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off.
  • Paid parental leave.
  • Relocation support as needed.
  • Visa sponsorship is available.
  • The role is based in San Francisco, California.

Tech Stack

Categories

AI & MLBackendData Engineering