XPENG

Staff Machine Learning Engineer - LLM Quantization & Deployment

XPENG
Apply
4 hours ago
Santa Clara, CA, USAStaff+

Base Salary

$215k - $364k/yr

Responsibilities

  • Develop VLA inference models and ensure numerical consistency with training models.
  • Productionize LLM quantization methods including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
  • Develop production-quality Python code with testing, observability, reproducibility, and failure handling.
  • Build model export, calibration, benchmarking, validation, and deployment pipelines.
  • Collaborate with VLA research to establish performance estimates and prove model feasibility.
  • Curate evaluation datasets and establish metrics to benchmark VLA performance.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Interface with field-testing and simulation teams for issue triage and autonomous-driving performance sign-off.
  • Collaborate with the in-vehicle software team on latency analysis and issue triage.
  • Collaborate with the training infrastructure team on QAT and model distillation.

Requirements

  • Master’s degree in computer science, computer engineering, or electrical engineering, or equivalent experience.
  • 3–5 years of industry experience.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization is preferred.
  • Experience with AWQ, GPTQ, SmoothQuant, or related methods is preferred.
  • Experience with LLM runtimes such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes is preferred.
  • Experience deploying LLMs on resource-constrained or heterogeneous hardware is preferred.
  • Contributions to model optimization, inference, compiler, or serving projects are preferred.
  • Publications at NeurIPS, ICML, ICLR, ACL, or related conferences are preferred.
  • Strong numerical analysis, systems engineering, communication, collaboration, problem-solving, and cross-functional working skills.

Benefits

  • Supportive and engaging work environment
  • Infrastructure and computational resources
  • Opportunity to work on cutting-edge AI and autonomous-driving technologies
  • Snacks, lunches, dinners, and fun activities

Tech Stack

Categories

XPENG

About XPENG

10,000+ employees

XPENG is a leading Chinese Smart EV company that designs, develops, manufactures, and markets Smart EVs that appeal to the large and growing base of technology-savvy middle-class consumers. Its mission is to drive Smart EV transformation with technology and data, shaping the mobility experience of the future. In order to optimize its customers’ mobility experience, XPeng develops in-house its full-stack advanced driver-assistance system technology and in-car intelligent operating system, as well as core vehicle systems including powertrain and the electrical/electronic architecture. XPeng is headquartered in Guangzhou, China. In 2021, the Company established its European headquarters in Amsterdam, along with other dedicated offices in Copenhagen, Munich, Oslo, and Stockholm.The Company’s Smart EVs are mainly manufactured at its plant in Zhaoqing and Guangzhou,Guangdong province. For more information, please visit https://www.xpeng.com/