18 hours ago
Shanghai, China or Beijing, ChinaIntern
Responsibilities
- Build and enhance high-performance LLM inference pipelines and optimize model execution, scalability, memory usage, and multi-GPU serving.
- Improve TensorRT compiler graph transformations, code generation, operator fusion, memory allocation, and optimization passes for NVIDIA GPUs.
- Design and tune CUDA kernels and DSL implementations for GEMM, MoE, Attention, Convolution, and other deep learning operations.
- Profile, analyze, and optimize GPU and deep learning software performance while collaborating with framework, research, CUDA, and hardware architecture teams.
Requirements
- Pursuing an M.S. or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Excellent problem-solving ability, curiosity about cutting-edge AI systems, and passion for GPU computing and deep learning software performance.
- TensorRT LLM track: strong Python programming, PyTorch experience, and understanding of inference and GPU acceleration.
- TensorRT Compiler track: proficiency in C++ and experience with compiler or performance optimization.
- CuTe DSL and CUDA Kernels track: skills in C/C++ and CUDA or parallel programming, familiarity with LLVM, MLIR, and compilers, and understanding of computer architecture and performance profiling, analysis, and optimization.
Benefits
- Internship opportunity with NVIDIA focused on AI computing platforms and GPU-accelerated software.
- Collaborative work with framework, research, CUDA, and hardware architecture teams.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
