18 hours ago
Shanghai, China or Beijing, ChinaIntern
Responsibilities
- Develop high-performance deep learning operators and fusion operators on NVIDIA GPUs.
- Optimize GPU kernels for cuBLAS, TensorRT, cuDNN, cuSparse, and cuTensor libraries.
- Analyze kernel performance on existing and new GPU architectures, identify bottlenecks, and propose improvements.
- Design and develop software infrastructure for kernel authoring and shipping.
- Apply emerging AI technologies to GPU kernel and related development workflows.
Requirements
- Pursuing a B.S., M.S., or PhD in computer science or a similar field.
- Strong programming skills in C/C++ and Python.
- Familiarity with GPU programming models and CUDA.
- Good understanding of compiler technologies and experience with LLVM and MLIR.
- Strong problem-solving, communication, and teamwork skills.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
