3 hours ago
Base Salary
$190k - $250k/yr
Responsibilities
- Design, implement, and optimize custom GPU kernels for large-scale AI systems.
- Profile ML operations, develop performance models, identify bottlenecks, and improve training and inference performance.
- Integrate low-level kernels into PyTorch, JAX, and internal runtimes.
- Collaborate with ML researchers, distributed systems engineers, model-serving teams, and NVIDIA/AMD hardware vendors.
- Contribute to tooling, documentation, benchmarking suites, and testing frameworks.
Requirements
- At least 5 years of industry or research experience in GPU kernel development or high-performance computing.
- Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Strong programming skills in C++ and Python and familiarity with ML frameworks.
- Deep expertise in CUDA/ROCm, GPU memory models, and performance optimization.
- Hands-on experience with Triton and/or JAX Pallas, PTX, GPU ASM, and custom GPU kernel development.
- Experience integrating low-level kernels into PyTorch, JAX, or similar frameworks.
- Experience with large-scale LLM training or inference workloads.
- Preferred: AMD GPU and ROCm optimization, JAX FFI, custom ML operators, vLLM, TensorRT, TPUs, XLA, or open-source ML systems, compilers, or GPU kernels.
Benefits
- Medical, dental, and vision insurance.
- 401k plan.
- Daily lunch, snacks, and beverages.
- Flexible time off.
- Competitive salary and equity.
