9 days ago
Base Salary
$190k - $250k/yr
Responsibilities
- Design, implement, and optimize custom GPU kernels using C++, PTX, CUDA, ROCm, Triton, and/or JAX Pallas.
- Profile and optimize ML operations for large-scale LLM training and inference.
- Integrate low-level GPU kernels into PyTorch, JAX, and custom internal runtimes.
- Develop performance models, identify bottlenecks, and deliver kernel-level improvements for AI workloads.
- Collaborate with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD hardware vendors.
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks for correctness and performance reproducibility.
Requirements
- At least 5 years of industry or research experience in GPU kernel development or high-performance computing.
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Strong programming skills in C++ and Python, with familiarity with ML frameworks.
- Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization strategies.
- Hands-on experience with Triton and/or JAX Pallas for custom kernel development.
- Strong understanding of PTX, GPU assembly, and low-level GPU execution.
- Extensive experience writing and optimizing custom GPU kernels in C++ and PTX.
- Proven ability to integrate low-level kernels into PyTorch, JAX, or similar frameworks.
- Experience with large-scale LLM training or inference.
- Preferred experience includes AMD GPUs, ROCm optimization, JAX FFI, custom ML operators, vLLM, TensorRT, TPUs, XLA, open-source ML systems, compilers, or GPU kernels.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
