5 hours ago
Base Salary
$190k - $250k/yr
Responsibilities
- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.
- Profile ML operations and optimize end-to-end performance for large-scale LLM training and inference.
- Integrate low-level GPU kernels into PyTorch, JAX, and custom internal runtimes.
- Develop performance models, identify bottlenecks, and deliver kernel-level performance improvements.
- Collaborate with ML researchers, distributed systems engineers, and model-serving teams.
- Work with NVIDIA and AMD hardware vendors on GPU architecture and compiler/toolchain capabilities.
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks for correctness and reproducibility.
Requirements
- At least 5 years of industry or research experience in GPU kernel development or high-performance computing.
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Strong programming skills in C++ and Python with familiarity with ML frameworks.
- Deep expertise in CUDA, ROCm, GPU memory models, and performance optimization.
- Hands-on experience with Triton and/or JAX Pallas for custom kernel development.
- Strong understanding of PTX, GPU ASM, and low-level GPU execution.
- Extensive experience writing and optimizing custom GPU kernels in C++ and PTX.
- Experience integrating low-level kernels into PyTorch, JAX, or similar frameworks.
- Experience with large-scale LLM training or inference.
- Preferred qualifications include AMD GPU and ROCm optimization, JAX FFI, custom ML operators, vLLM, TensorRT, TPUs, XLA, accelerator programming, and contributions to open-source ML systems, compilers, or GPU kernels.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
