20 hours ago
Shanghai, China or Beijing, ChinaIntern
Responsibilities
- Create and maintain SKILL, Wiki, and agent harness tools.
- Develop TileGym, the Triton CUDA TileIR backend, and CUDA Tile.
- Develop highly optimized deep learning kernels using a tile-based GPU programming model.
- Perform end-to-end performance optimization, analysis, and tuning.
- Conduct performance modeling, profiling, debugging, and code optimization.
Requirements
- Pursue a university degree in engineering, computer science, or a related field; master's or doctoral study is preferred.
- Understand core agentic-system components, including LLM APIs, prompting, tool use/function calling, agent loops, planning, reasoning, memory, RAG, skills, MCP, sub-agents, multi-agent architectures, context engineering, and harness engineering.
- Have excellent C/C++ programming and software design skills.
- Have performance modeling, profiling, debugging, code optimization, or CPU/GPU architectural knowledge.
- Python experience is a plus.
- MLIR experience is a plus.
- GPU programming experience with CUDA or OpenCL is desired.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
