1 day ago
Base Salary
$184k - $357k/yr
Responsibilities
- Optimize, analyze, tune, and scale deep learning models for LLM, multimodal, and generative AI workloads.
- Develop and contribute features and code to vLLM, SGLang, FlashInfer, NVIDIA inference libraries, and related deep learning software.
- Implement current algorithms from the deep learning community for public release in inference frameworks.
- Improve model-serving performance across different NVIDIA GPU and accelerator architectures.
- Collaborate with teams across frameworks, NVIDIA libraries, and inference optimization initiatives.
Requirements
- Master’s degree, PhD, or equivalent experience in Computer Engineering, Computer Science, EECS, AI, or a related field.
- 6+ years of relevant software development experience.
- Excellent C/C++ programming and software design skills.
- Python experience is a plus, and software Agile skills are helpful.
- Experience training, deploying, or optimizing deep learning model inference in production is a plus.
- Experience with performance modeling, profiling, debugging, code optimization, or CPU/GPU architecture is a plus.
- Preferred experience contributing to PyTorch, vLLM, or SGLang; working with NCCL or NVSHMEM; shipping enterprise products; or programming GPUs with CUDA, OAI Triton, or CUTLASS.
Benefits
- Eligible for equity and benefits.
- Base salary ranges by level are $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5.
- Applications will be accepted at least until September 14, 2026.
- This posting is for an existing vacancy.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
