18 hours ago
Shanghai, China or Beijing, ChinaIntern
Responsibilities
- Write highly tuned compute kernels for core deep learning operations including matrix multiplication, mixture-of-experts, and attention.
- Optimize linear algebra and deep learning operations for NVIDIA GPUs and future GPU architectures.
- Perform debugging, performance analysis, bottleneck identification, resource optimization, and throughput improvement.
- Support regression testing and CI/CD flows.
- Collaborate with compiler, deep learning performance, training and inference, hardware, and architecture teams.
Requirements
- Pursuing a master’s or PhD in computer science, computer engineering, applied mathematics, or a related field.
- Strong programming and software design skills, including debugging, performance analysis, and test design.
- Experience with performance-oriented parallel programming, including OpenMP or pthreads, even if not on GPUs.
- Solid understanding of computer architecture and some assembly programming.
- Preferred experience tuning deep learning library kernel code, CUDA GPU programming, numerical methods, linear algebra, LLVM, TVM tensor expressions, or TensorFlow MLIR.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
