7 days ago
Shanghai, China or Beijing, ChinaMid Level
H1B Sponsor
Responsibilities
- Architect end-to-end LLM pretraining, fine-tuning, inference, RAG, and agentic inference solutions using NVIDIA platforms.
- Collaborate with enterprise customers to understand business challenges and design tailored solutions.
- Lead LLM training, distributed optimization, and performance tuning for throughput, latency, and memory efficiency.
- Design and integrate RAG workflows and agentic inference pipelines into customer systems.
- Collaborate with NVIDIA engineering teams and support pre-sales workshops, demonstrations, and technical activities.
Requirements
- Master’s or Ph.D. in Computer Science, Artificial Intelligence, or equivalent experience.
- At least 4 years of hands-on AI experience focused on open-source LLM training, fine-tuning, and production inference optimization.
- Strong understanding of mainstream LLM architectures and proficiency with PyTorch and Hugging Face Transformers.
- Knowledge of GPU computing, cluster architecture, and distributed parallel LLM training and inference.
- Competency in agentic inference design and applying AI agents to business challenges.
- Strong communication skills for explaining complex technical concepts to technical and non-technical stakeholders.
- Preferred experience with TRT-LLM, Megatron-LM, NVIDIA NeMo, quantization, KV Cache tuning, memory optimization, Docker, Kubernetes, multi-GPU parallelism, and large-scale GPU cluster management.
Benefits
- Competitive salary and a generous benefits package.
Tech Stack
Categories
Solutions Engineering
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
