XPENG

Staff Machine Learning Engineer - LLM Quantization & Deployment

XPENG
Apply
26 days ago
Santa Clara, CA, USAStaff+
H1B sponsor

Base Salary

$215k - $364k/yr

Responsibilities

  • Develop VLA inference models and ensure numerical consistency with training models.
  • Productionize LLM quantization methods including PTQ, QAT, mixed-precision inference, INT8, FP4, and lower-bit techniques.
  • Develop production-quality Python code with testing, observability, reproducibility, and failure handling.
  • Build model export, calibration, benchmarking, validation, and deployment pipelines.
  • Collaborate with VLA research to establish performance estimates and prove model feasibility.
  • Curate evaluation datasets and establish metrics to benchmark VLA performance.
  • Analyze numerical errors, accuracy regressions, and performance trade-offs.
  • Develop PTQ and QAT orchestration workflows.
  • Interface with field-testing and simulation teams for issue triage and autonomous-driving performance sign-off.
  • Collaborate with the in-vehicle software team on latency analysis and issue triage.
  • Collaborate with the training infrastructure team on QAT and model distillation.

Requirements

  • Master’s degree in computer science, computer engineering, or electrical engineering, or equivalent experience.
  • 3–5 years of industry experience.
  • Strong understanding of Transformer architectures and LLM inference.
  • Hands-on experience quantizing or deploying deep learning models in production.
  • Proficiency with PyTorch and at least one inference or compilation stack.
  • Strong Python programming and software engineering skills.
  • Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization is preferred.
  • Experience with AWQ, GPTQ, SmoothQuant, or related methods is preferred.
  • Experience with LLM runtimes such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes is preferred.
  • Experience deploying LLMs on resource-constrained or heterogeneous hardware is preferred.
  • Contributions to model optimization, inference, compiler, or serving projects are preferred.
  • Publications at NeurIPS, ICML, ICLR, ACL, or related conferences are preferred.
  • Strong numerical analysis, systems engineering, communication, collaboration, problem-solving, and cross-functional working skills.

Benefits

  • Supportive and engaging work environment
  • Infrastructure and computational resources
  • Opportunity to work on cutting-edge AI and autonomous-driving technologies
  • Snacks, lunches, dinners, and fun activities

Tech Stack

Categories

XPENG

About XPENG

1,001-5,000 employees

XPENG designs, manufactures, and sells smart electric vehicles for consumers, integrating in-house advanced driver-assistance systems and connected in-car software. Founded in 2014 and headquartered in Guangzhou, it operates plants in Zhaoqing and Guangzhou, maintains a European HQ in Amsterdam, and develops eVTOL aircraft via XPENG AEROHT.

Contact me