7 days ago
Responsibilities
- Partner with Sales, BD, and CPM teams to drive adoption of NVIDIA GPU and AI infrastructure technologies among major Chinese CSP accounts.
- Provide end-to-end technical consultation on GPU cluster architecture, AI workload deployment, heterogeneous computing, and software-stack optimization.
- Optimize CPU-GPU co-processing, data movement, latency, and throughput for reinforcement learning and Agentic AI workloads.
- Analyze GPU workload bottlenecks and implement system-, kernel-, and framework-level tuning for training, inference, reinforcement learning, and gaming workloads.
- Lead open-source architecture contributions, upstream optimized patches, and China-localized AI infrastructure best practices.
- Act as the technical liaison between Chinese CSP customers and NVIDIA global engineering, product, and R&D teams.
- Lead technical workshops, training, PoCs, production pilots, reference designs, and large-scale deployment projects.
- Monitor AI infrastructure, LLM inference, cloud gaming, Agentic AI, and data center architecture trends to inform strategy.
- Mentor junior solutions architects and standardize technical engagement and delivery practices.
Requirements
- Bachelor’s, master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field; equivalent industry experience is valued.
- 7+ years of hands-on experience in GPU architecture, AI system optimization, large-scale data center infrastructure, or hyperscale cloud computing.
- Experience with AI training and inference, distributed computing, or HPC workloads.
- Deep knowledge of GPU microarchitecture, CUDA, GPU memory hierarchy, scheduling, profiling, bottleneck analysis, and end-to-end workload tuning.
- Strong C/C++ and Python programming skills, with familiarity with CUDA kernels, compiler toolchains, PyTorch, TensorRT, and distributed system tuning.
- Hands-on experience with major Chinese CSPs or global hyperscalers and their public-cloud AI service architectures and workloads.
- Strong technical communication, presentation, cross-functional collaboration, project ownership, and independent delivery capabilities.
- Familiarity with NVIDIA GPU data center hardware, TensorRT-LLM, Dynamo, NCCL, and the CUDA software stack is a significant plus.
- Experience with cluster networking diagnostics, Linux kernel, drivers, virtualization, AI infrastructure open source, GPU optimization libraries, or distributed computing frameworks is advantageous.
- Experience with Vera or Grace CPU-GPU co-optimization, AI agents, reinforcement-learning post-training, or long-context LLM optimization is advantageous.
Benefits
- Competitive salary and a generous benefits package.
- Opportunity to work with NVIDIA engineering teams and major CSP customers on large-scale AI infrastructure.
- The posting describes a rapidly growing engineering organization and work requiring autonomy and technical impact.
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
