4 hours ago
Base Salary
$190k - $250k/yr
Responsibilities
- Design, implement, and optimize custom GPU kernels for large-scale AI workloads.
- Profile end-to-end ML operation performance, identify bottlenecks, and deliver kernel-level improvements.
- Integrate low-level kernels into PyTorch, JAX, and internal runtimes.
- Develop performance models and benchmarking, tooling, documentation, and testing resources.
- Collaborate with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD hardware vendors.
Requirements
- 5+ years of industry or research experience in GPU kernel development or high-performance computing.
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Strong programming skills in C++ and Python, with familiarity with ML frameworks.
- Deep expertise in CUDA/ROCm, GPU memory models, GPU ASM, PTX, and performance optimization.
- Hands-on experience with Triton and/or JAX Pallas and integrating kernels into PyTorch, JAX, or similar frameworks.
- Experience with large-scale LLM training or inference.
- Preferred qualifications include AMD GPU and ROCm optimization, JAX FFI, custom ML operators, vLLM, TensorRT, TPUs, XLA, compilers, and open-source ML systems.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
