1 day ago
Sunnyvale, CA, USAStaff+
Base Salary
$336k - $359k/yr
Responsibilities
- Profile large-scale ML workloads to identify performance bottlenecks using tools such as NVIDIA Nsight Systems.
- Design and implement efficiency improvements involving parallelism, model compilation, and mixed precision to maximize MFU and throughput.
- Build observability tools to track MFU, throughput, latency, and other performance indicators.
- Develop benchmarking tools to measure efficiency improvements and regressions.
- Collaborate with research teams to integrate training-efficiency improvements and promote performance optimization practices.
Requirements
- 10+ years of industry experience driving performance engineering across ML systems, GPU compute infrastructure, distributed platforms, or a similar field.
- Experience optimizing large-scale jobs on GPU compute clusters.
- Experience working with platform teams and research teams.
- Experience writing, reporting, and tracking performance benchmarks in an open and accessible way.
- Ability to write high-quality, well-structured, and tested Python code.
- Bachelor’s or master’s degree in machine learning, computer science, engineering, or a related technical discipline, or equivalent experience.
- Desirable experience with concurrent, parallel, and distributed computing.
- Desirable experience using NVIDIA Nsight Systems or other system profilers.
- Desirable experience implementing GPU kernels with CUDA, Triton, or similar technologies.
- Knowledge of computing fundamentals related to code performance, security, and reliability.
Benefits
- Full-time employment based in Sunnyvale, California, with a hybrid work arrangement.
- Competitive equity package.
- Inclusive interview experience with accommodations or adjustments available upon request.
