
Research Engineer, Infrastructure, Inference
Thinking Machines Lab2 months ago
Base Salary
$350k - $475k/yr
Responsibilities
- Work with researchers and engineers to bring cutting-edge AI models into production.
- Enable high-performance inference for novel model architectures.
- Design and implement techniques, tools, and architectures that improve performance, latency, throughput, and efficiency.
- Optimize code and GPU compute fleets for hardware utilization, bandwidth, and memory efficiency.
- Extend orchestration frameworks for distributed inference, evaluation, and large-batch serving.
- Establish reliability, observability, and reproducibility standards across the inference stack.
- Share technical learnings through internal documentation, open-source libraries, and technical reports.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, or a similar field.
- Understanding of deep learning frameworks and their underlying system architectures.
- Experience with inference serving systems optimized for throughput and latency.
- Strong engineering skills and the ability to write performant, maintainable code and debug complex codebases.
- Ability to collaborate across teams and take initiative to deliver work.
- Preferred: experience training or supporting large-scale language models with hundreds of billions of parameters or more.
- Preferred: understanding of distributed compute systems, GPU parallelism, and hardware-aware optimizations.
- Preferred: contributions to open-source ML or systems infrastructure projects.
- Preferred: a track record of improving research productivity through infrastructure or process improvements.
Benefits
- Generous health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- This is an evergreen, ongoing role with applications reviewed continuously.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.