
Research Engineer, Infrastructure, Inference
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Bring cutting-edge AI models into production alongside researchers and engineers.
- Enable high-performance inference for novel model architectures.
- Design and implement techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
- Optimize the codebase and compute fleet, including GPUs, for hardware utilization.
- Extend orchestration frameworks for distributed inference, evaluation, and large-batch serving.
- Establish reliability, observability, and reproducibility standards across the inference stack.
- Share learnings through documentation, open-source libraries, and technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
- Understanding of deep learning frameworks and their underlying system architectures.
- Experience with inference serving systems optimized for throughput and latency.
- Strong engineering skills, including the ability to write performant, maintainable code and debug complex codebases.
- Preferred experience training or supporting language models with hundreds of billions of parameters or more.
- Preferred understanding of distributed compute systems, GPU parallelism, and hardware-aware optimization.
- Open-source contributions to ML or systems infrastructure projects are preferred.
- A track record of improving research productivity through infrastructure design or process improvements is preferred.
Benefits
- Health, dental, and vision benefits
- Unlimited paid time off
- Paid parental leave
- Relocation support as needed
- Visa sponsorship is available
- Role is based in San Francisco, California
- This is an evergreen role reviewed on an ongoing basis; applicants should not reapply more than once every six months