1 day ago
Taipei, TaiwanIntern
Responsibilities
- Build and enhance high-performance LLM inference pipelines.
- Analyze and optimize LLM execution, scalability, and memory usage.
- Deliver efficient multi-GPU model serving for Agentic AI workloads.
- Improve TensorRT compiler graph transformations and code generation for NVIDIA GPUs.
- Develop compiler optimization passes, refine operator fusion, and optimize memory usage.
- Collaborate with framework, research, CUDA, and hardware architecture teams to accelerate deep learning inference.
Requirements
- Pursuing a B.S., M.S., Ph.D., or equivalent in Computer Science, Computer Engineering, Electrical Engineering, Applied Mathematics, or a related field.
- Demonstrated problem-solving ability, curiosity about modern AI systems, and interest in GPU computing and deep learning software performance.
- For the TensorRT-LLM track, strong Python programming, PyTorch experience, and understanding of LLM inference operations and GPU acceleration.
- For the TensorRT Compiler track, C++ proficiency and experience with compilers or performance optimization.
Benefits
- Internship opportunity with NVIDIA’s AI Compute team in Taiwan.
- Equal opportunity employer with reasonable accommodation available for individuals with disabilities.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
