Nvidia

Principal Architect, AI Inference Networking - NIXL and Dynamo

Nvidia
Apply
18 hours ago
Shanghai, China or Shenzhen, ChinaStaff+

Responsibilities

  • Drive NIXL adoption in China from architecture discussions through prototypes, benchmarks, production rollout, and upstream contributions.
  • Profile customer inference stacks and optimize KV-cache movement, memory registration, transfer scheduling, throughput, and time to first token.
  • Write NIXL backends and plugins, extend Dynamo integrations, and contribute optimizations to NIXL and GPUDirect-class data paths.
  • Define dynamic data-path APIs supporting elastic scaling, worker churn, runtime rerouting, and production inference clusters.
  • Represent regional requirements in NVIDIA’s roadmap and influence its technical direction.
  • Build the local NIXL community through talks, benchmarks, and reference architectures.
  • Collaborate with NIXL and Dynamo engineers across China, Israel, and the United States.

Requirements

  • M.Sc. or Ph.D. in computer science, electrical engineering, or computer engineering, or equivalent depth earned in industry.
  • 12+ years building or optimizing large-scale distributed systems, AI inference or training infrastructure, communication libraries, high-performance networking, or HPC runtimes.
  • Strong systems programming skills in C++ and Python, including shipped performance-critical code.
  • Hands-on experience with RDMA networking using InfiniBand or RoCE and GPU memory movement.
  • Working knowledge of modern LLM serving, including prefill/decode disaggregation, KV-cache management, tensor parallelism, and pipeline parallelism.
  • Experience with at least one serving stack: Dynamo, TensorRT-LLM, vLLM, SGLang, or Triton.
  • Strong communication skills and technical credibility with customer engineering teams and core library maintainers.
  • Fluent Mandarin and professional English, with comfort working across China, Israel, and US time zones.
  • Preferred qualifications include open-source contributions to NIXL, Dynamo, NCCL, vLLM, SGLang, TensorRT-LLM, or Mooncake; inference deployments at thousands-of-GPUs scale; CUDA and kernel or driver-level code experience; and a history of building an adopted technical community or project.

Tech Stack

Categories

BackendSolutions Engineering
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me