10 months ago
Palo Alto, CA, USASenior
Responsibilities
- Develop, integrate, and optimize CUDA kernels for large-scale AI training, inference, and reinforcement learning workloads.
- Profile and optimize GPU performance for custom compute-bound and memory-bound workloads.
- Integrate custom kernels into training and inference frameworks such as PyTorch, Megatron, vLLM, and TorchTitan.
- Build performance tools, benchmarks, integration layers, and GPU-accelerated primitives for graph reasoning, symbolic computation, and hardware simulation.
- Collaborate with AI researchers and semiconductor experts to translate domain-specific workloads into high-performance GPU code.
- Release kernels and tooling as contributions to open-source AI and HPC ecosystems.
Requirements
- Experience writing and optimizing CUDA kernels for large-scale AI workloads, including attention, routing, graph-based operations, or physics-inspired operators.
- Experience profiling and optimizing GPU performance.
- Experience integrating custom kernels with AI training and inference frameworks.
- Experience with NVIDIA hardware and software stacks, including Hopper, Blackwell, NVLink, NCCL, and Triton.
- Experience building GPU-accelerated primitives for graph reasoning, symbolic computation, or hardware simulation.
- Ability to collaborate with AI researchers and semiconductor experts on domain-specific performance engineering.
