NIO Inc.

AI Infrastructure Engineer

NIO Inc.
Apply
3 months ago
San Jose, CA, USASenior

Base Salary

$192k - $250k/yr

Responsibilities

  • Design and implement high-performance, scalable inference systems for LLMs and VLMs across cloud, edge, and hybrid edge-cloud platforms.
  • Develop and optimize custom kernels and operators for GPU, NPU, DSP, and other hardware accelerators.
  • Integrate KV-cache management, tensor and model parallelism, quantization, and memory-efficient execution into production inference systems.
  • Partner with systems and hardware teams to optimize hardware-software integration across diverse compute environments.
  • Translate architectural requirements into robust, maintainable software meeting performance, safety, and reliability standards.
  • Define and drive the roadmap for LLM/VLM inference in the AIOS stack.
  • Monitor industry and competitor developments and apply relevant AI and large-scale systems engineering practices.

Requirements

  • At least 5 years of hands-on software development experience building and optimizing AI inference systems at scale.
  • Direct experience with LLM/VLM model internals, Transformer architectures, inference bottlenecks, and optimization techniques.
  • Strong expertise in kernel development, parallelism, memory optimization, and distributed inference systems.
  • Proficiency with GPU/NPU programming, CUDA or vendor-specific SDKs, compiler toolchains, and PyTorch or TensorFlow.
  • Strong C/C++ programming skills and experience delivering high-performance production software.
  • Strong foundation in computer architecture, systems programming, CPU/GPU pipelines, memory hierarchy, scheduling, and embedded systems.
  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related technical field.
  • Master's or PhD degree and 5 years of industry experience are preferred.
  • Experience with inference serving systems for large models, including batching, scheduling, caching, and load balancing, is preferred.
  • Expertise in hardware-aware optimization such as kernel fusion, mixed precision, quantization, and pruning is preferred.
  • Familiarity with edge and embedded AI, real-time constraints, and limited-resource optimization is preferred.
  • Contributions to widely used AI frameworks, libraries, or performance-critical software are preferred.
  • Strong communication and cross-functional collaboration skills.

Benefits

  • Full-time employees are eligible for medical plans including Anthem Blue Cross, HSA, and Kaiser HMO, with $0 employee-only coverage.
  • Dental and vision plans offer options with $0 paycheck contribution for employees and eligible dependents.
  • Company-paid HSA contributions are available with the High Deductible Anthem Blue Cross plan.
  • Healthcare and dependent care FSAs, 401(k) with BrokerageLink, company-paid life and disability insurance, and an Employee Assistance Program are provided.
  • Benefits include sick and vacation time, 13 paid holidays, paid parental leave, and paid disability leave subject to eligibility periods.
  • Additional benefits include voluntary life and AD&D insurance, pet insurance, commuter benefits, mobile phone credit, free lunch and snacks, an onsite gym, and employee discounts and perks.

Tech Stack

CC++PyTorchTensorFlow
NIO Inc.

About NIO Inc.

5,001-10,000 employees
Contact me