2 days ago
Base Salary
$184k - $357k/yr
Responsibilities
- Lead end-to-end performance analysis of LLM and VLM inference workloads and optimize latency, throughput, efficiency, KV-cache capacity, and scale.
- Build speed-of-light and roofline models and use profiler data to identify and eliminate host, kernel, memory, communication, and scheduling bottlenecks.
- Tune batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, model parallelism, and other serving techniques.
- Develop and optimize attention, matrix multiplication, mixture-of-experts, quantization, and data-movement kernels using CUDA, CUTLASS, Triton, or related technologies.
- Establish repeatable benchmarks, canonical run records, and performance-regression gates across models, hardware, topology, software, precision, and workloads.
- Collaborate with model, framework, kernel, networking, and GPU architecture teams and contribute improvements to TensorRT-LLM, vLLM, SGLang, or related projects.
Requirements
- More than 6 years of experience in full-stack LLM/VLM inference performance across models, serving, distributed runtimes, kernels, and hardware.
- Strong programming skills in Python, Rust and/or C++, with hands-on experience in CUDA or another GPU programming environment.
- Expertise in speed-of-light analysis, roofline models, microbenchmarks, NVIDIA Nsight Systems, and NVIDIA Nsight Compute.
- Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats.
- Practical experience optimizing inference servers and model execution using batching, scheduling, KV-cache management, quantization, speculative decoding, and parallelism strategies.
- Understanding of distributed systems and networking for accelerated computing, including collectives, topology, and scale-up versus scale-out performance.
- BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- Preferred qualifications include contributions to TensorRT-LLM, vLLM, SGLang, PyTorch, CUDA, Triton, or NCCL; AI-agent-supported performance workflows; published research or technical presentations; and experience with new LLM/VLM architectures, long-context inference, mixture-of-experts models, multimodal pipelines, or large-scale distributed serving.
Benefits
- Competitive salary, equity, and a generous benefits package are offered.
- The posting is for an existing vacancy, with applications accepted at least until October 2, 2026.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
