Nvidia

Senior Performance Software Engineer, Deep Learning Libraries

Nvidia
Apply
19 hours ago
Shanghai, China or Beijing, ChinaSenior

Responsibilities

  • Write highly tuned compute kernels for core deep learning operations such as matrix multiplication, convolutions, and normalizations.
  • Optimize linear algebra and deep learning operations for NVIDIA GPUs and future GPU architectures.
  • Analyze performance, identify bottlenecks, optimize resource utilization, and improve throughput.
  • Apply software engineering practices including debugging, test design, regression testing, and CI/CD flows.
  • Collaborate with CUDA compiler, deep learning performance, training and inference, hardware, and architecture teams.

Requirements

  • Master’s or PhD degree in Computer Science, Computer Engineering, Applied Math, or a related field, or equivalent experience.
  • At least 2 years of relevant industry experience.
  • Strong C++ programming and software design skills, including debugging, performance analysis, and test design.
  • Experience with performance-oriented parallel programming, such as OpenMP or pthreads.
  • Solid understanding of computer architecture and some assembly programming experience.

Tech Stack

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me