1 day ago
Shanghai, China or Beijing, ChinaMid Level
Responsibilities
- Drive implementation, deployment, and optimization of NVIDIA Inference Microservices solutions for enterprise and industry AI workloads.
- Package and serve open-source, NVIDIA, and customer-proprietary models through standardized, containerized APIs across on-premises, cloud, and hybrid environments.
- Optimize high-volume inference and rollout workloads for LLMs and VLMs.
- Evaluate and tune NIM models, including performance, memory, and model-quality optimization.
- Deliver technical projects, demonstrations, and client support tasks.
- Provide technical support and guidance to customers to facilitate adoption and implementation of NVIDIA technologies and products.
- Collaborate with research, engineering, infrastructure, product, and customer-facing teams to expand the AI solutions portfolio.
Requirements
- Master’s degree or higher in Computer Science, Machine Learning, Electrical Engineering, Mathematics, or a related technical field, or equivalent experience.
- At least 2 years of hands-on experience in machine learning engineering, applied research, LLM/VLM inference, or reinforcement-learning rollout.
- Production-quality Python and PyTorch skills, including distributed GPU training, solution profiling, debugging, and memory optimization.
- Working knowledge of transformer architectures, performance optimization, rollout sampling strategies, structured generation, and model-quality evaluation.
- Strong written and verbal communication and effective cross-functional collaboration skills.
- Preferred qualifications include publications, open-source contributions, significant technical projects, LLM/VLM or agent-system experience, and familiarity with enterprise AI deployment or foundation-model adaptation.
- Preferred experience includes programmatic verification, simulators, compilers, execution sandboxes, APIs, external reward tools, SLIME, or NeMo-RL.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
