XPENG

Senior Machine Learning Engineer - LLM Quantization & Deployment

XPENG
Apply
4 hours ago
Santa Clara, CA, USASenior

Base Salary

$175k - $296k/yr

Responsibilities

  • Develop VLA inference models and ensure numerical consistency with training models.
  • Productionize LLM quantization methods including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
  • Develop production-quality Python code with testing, observability, reproducibility, and failure handling.
  • Build model export, calibration, benchmarking, validation, and deployment pipelines.
  • Collaborate with VLA model research teams to estimate performance and prove model feasibility.
  • Curate evaluation datasets and establish metrics for systematic VLA benchmarking.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Interface with field-testing and simulation teams for issue triage and autonomous driving performance sign-off.
  • Collaborate with in-vehicle software teams on latency analysis and issue triage.
  • Work with the training infrastructure team on QAT and model distillation.

Requirements

  • Master’s degree in computer science, computer engineering, or electrical engineering, or equivalent experience, with 1–3 years of industry experience; new graduates are welcome.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization is preferred.
  • Experience with AWQ, GPTQ, SmoothQuant, or related methods is preferred.
  • Experience with LLM runtimes such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes is preferred.
  • Experience deploying LLMs on resource-constrained or heterogeneous hardware is preferred.
  • Contributions to model optimization, inference, compiler, or serving projects are preferred.
  • Publications at NeurIPS, ICML, ICLR, ACL, or related conferences are preferred.

Benefits

  • Supportive and engaging work environment
  • Infrastructure and computational resources to support the work
  • Opportunity to work with cutting-edge technologies and leading researchers
  • Snacks, lunches, dinners, and fun activities
  • Bonus, equity, and benefits in addition to base salary

Tech Stack

Categories

XPENG

About XPENG

10,000+ employees

XPENG is a leading Chinese Smart EV company that designs, develops, manufactures, and markets Smart EVs that appeal to the large and growing base of technology-savvy middle-class consumers. Its mission is to drive Smart EV transformation with technology and data, shaping the mobility experience of the future. In order to optimize its customers’ mobility experience, XPeng develops in-house its full-stack advanced driver-assistance system technology and in-car intelligent operating system, as well as core vehicle systems including powertrain and the electrical/electronic architecture. XPeng is headquartered in Guangzhou, China. In 2021, the Company established its European headquarters in Amsterdam, along with other dedicated offices in Copenhagen, Munich, Oslo, and Stockholm.The Company’s Smart EVs are mainly manufactured at its plant in Zhaoqing and Guangzhou,Guangdong province. For more information, please visit https://www.xpeng.com/