23 days ago
San Jose, CA, USASenior
Base Salary
$200k - $400k/yr
Responsibilities
- Optimize training performance for 100B+ parameter models across 100,000+ GPUs.
- Advise on accelerator selection, cluster topology, scheduling, networking, and hardware procurement for future scaling.
- Write and optimize custom Triton and CUDA kernels.
- Build tooling and dashboards for performance monitoring, regression detection, and root-cause analysis.
- Optimize data loading and preprocessing pipelines to prevent I/O bottlenecks.
- Improve checkpointing, fault tolerance, and elastic restart for large-scale training jobs.
- Co-design model architectures and training recipes with researchers, including activation checkpointing, mixed precision, and sequence packing strategies.
- Extend kernel compilers such as Triton and Gluon for faster iteration and support for custom and non-NVIDIA accelerators.
- Build agentic systems that generate, benchmark, and iterate on custom kernels.
- Evaluate emerging accelerator architectures and lead proof-of-concept ports and benchmarks.
- Evaluate model and data parallelism strategies including FSDP, context parallelism, and expert parallelism.
Requirements
- Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- 3+ years of experience in AI performance engineering, including significant experience leading large-scale performance improvement projects.
- Deep understanding of GPU architecture and performance characteristics, including memory bandwidth, compute-bound versus memory-bound operations, and occupancy.
- Proficiency with profiling tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, HTA, or similar tools.
- Understanding of collective communication and modern networking concepts including NCCL, RDMA, NVLink, InfiniBand/RoCE, and topology-aware placement.
- Strong Python and CUDA/C++ skills, with the ability to read and modify framework internals.
- Experience debugging performance regressions and instability at scale, including stragglers, hangs, out-of-memory issues, and numerical divergence.
- Experience defining and using hardware-efficiency metrics such as MFU and HFU.
- Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration is a bonus.
- Open-source contributions to ML systems projects such as PyTorch, Megatron-LM, vLLM, DeepSpeed, or JAX are a bonus.
- Exposure to non-NVIDIA accelerators and heterogeneous fleet management is a bonus.
Benefits
- This is a full-time position.
- Additional compensation components and benefits may be provided depending on the specific role and shared if an employment offer is extended.
