GrepJob
Thinking Machines Lab

Research Engineer, Infrastructure, Inference

Thinking Machines Lab
Apply
5 days ago

Base Salary

$350k - $475k/yr

Responsibilities

  • Bring cutting-edge AI models into production alongside researchers and engineers.
  • Enable high-performance inference for novel model architectures.
  • Design and implement techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
  • Optimize the codebase and compute fleet, including GPUs, for hardware utilization.
  • Extend orchestration frameworks for distributed inference, evaluation, and large-batch serving.
  • Establish reliability, observability, and reproducibility standards across the inference stack.
  • Share learnings through documentation, open-source libraries, and technical reports.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
  • Understanding of deep learning frameworks and their underlying system architectures.
  • Experience with inference serving systems optimized for throughput and latency.
  • Strong engineering skills, including the ability to write performant, maintainable code and debug complex codebases.
  • Preferred experience training or supporting language models with hundreds of billions of parameters or more.
  • Preferred understanding of distributed compute systems, GPU parallelism, and hardware-aware optimization.
  • Open-source contributions to ML or systems infrastructure projects are preferred.
  • A track record of improving research productivity through infrastructure design or process improvements is preferred.

Benefits

  • Health, dental, and vision benefits
  • Unlimited paid time off
  • Paid parental leave
  • Relocation support as needed
  • Visa sponsorship is available
  • Role is based in San Francisco, California
  • This is an evergreen role reviewed on an ongoing basis; applicants should not reapply more than once every six months