
Research Engineer, Infrastructure, Kernels
Thinking Machines Lab2 months ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design and implement custom ML kernels for core language model operations, optimized for modern GPU and accelerator architectures.
- Develop compute primitives that reduce memory bandwidth bottlenecks and improve kernel efficiency.
- Collaborate with research teams to align kernel optimizations with model architecture and algorithmic objectives.
- Develop and maintain reusable kernel libraries and performance benchmarks for internal model training.
- Improve infrastructure stability, scalability, reproducibility, precision consistency, and compute utilization.
- Document and share technical insights through internal talks, technical papers, or open-source contributions.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
- Strong engineering skills and the ability to write performant, maintainable code and debug complex codebases.
- Understanding of deep learning frameworks and their underlying system architectures, including PyTorch or JAX.
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
- Demonstrated ability to analyze, profile, and optimize compute-intensive workloads.
- Experience training or supporting language models with tens of billions of parameters or more is preferred.
- Experience improving research productivity through infrastructure or process improvements is preferred.
- Experience developing or tuning kernels for PyTorch, JAX, or custom accelerators is preferred.
- Familiarity with tensor parallelism, pipeline parallelism, or distributed data processing frameworks is preferred.
- Experience with low-precision formats such as FP8, INT8, or block floating point and compiler stacks such as XLA or TVM is preferred.
- Contributions to open-source GPU, ML systems, or compiler optimization projects are preferred.
- Prior research or engineering experience in numerical optimization, communication-efficient training, or scalable AI infrastructure is preferred.
Benefits
- Health, dental, and vision benefits.
- Unlimited paid time off.
- Paid parental leave.
- Relocation support as needed.
- Visa sponsorship is available.
- The role is based in San Francisco, California.
- The position is an evergreen role reviewed on an ongoing basis, with applicants encouraged to reapply no more than once every six months.
Tech Stack
Categories
About Thinking Machines Lab
Thinking Machines Lab develops AI and generative AI software and conducts applied research to help organizations make data-driven decisions. The company builds products and data science solutions for enterprise use cases, pairing foundational models with practical tooling and services across industries. It is privately held and headquartered in San Francisco.