
Research Engineer, Infrastructure, Kernels
Thinking Machines Lab5 days ago
Base Salary
$350k - $475k/yr
Responsibilities
- Design and implement custom ML kernels for attention, matrix multiplication, gating, normalization, and other core LLM operations.
- Develop compute primitives that reduce memory bandwidth bottlenecks and improve kernel efficiency on modern GPU and accelerator architectures.
- Collaborate with research teams to align kernel optimizations with model architecture and algorithmic goals.
- Build and maintain reusable kernel libraries and performance benchmarks for internal model training.
- Improve infrastructure stability, scalability, reproducibility, precision consistency, and compute utilization.
- Share technical insights through internal talks, technical papers, or open-source contributions.
Requirements
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or a similar field.
- Strong engineering skills with the ability to write performant, maintainable code and debug complex codebases.
- Understanding of deep learning frameworks such as PyTorch and JAX and their underlying system architectures.
- Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.
- Demonstrated ability to analyze, profile, and optimize compute-intensive workloads.
- Experience with large-scale language model training, distributed parallelism, low-precision formats, compiler stacks, numerical optimization, scalable AI infrastructure, or related open-source projects is preferred.
Benefits
- Health, dental, and vision benefits
- Unlimited paid time off
- Paid parental leave
- Relocation support as needed
- Visa sponsorship is available
- Role is based in San Francisco, California
- Evergreen role reviewed on an ongoing basis; applicants should not reapply more than once every 6 months