5 months ago
Base Salary
$244k - $413k/yr
Responsibilities
- Lead the design of end-to-end Vision-Language-Action architectures connecting multimodal perception, linguistic reasoning, and action generation
- Drive research and development of generative world models and latent dynamics for controllable driving simulation
- Apply online and offline reinforcement learning and imitation learning to improve long-horizon driving policies and multi-agent interaction handling
- Define scaling laws and oversee data curation, automated labeling, and post-training for multi-billion-parameter driving foundation models
- Lead model adaptation for overseas road conditions, traffic laws, and driving cultures
Requirements
- 5–8 years of deep learning experience
- Significant experience with VLM, VLA, or embodied AI
- Experience training and deploying foundation models, including Transformers and LLMs, at scale
- Deep understanding of sequential decision-making, world models, or policy-gradient methods
- Mastery of PyTorch and distributed training technologies such as DeepSpeed and Megatron
- Ability to balance advanced research with the deterministic requirements of L4 production vehicles
Benefits
- Supportive and engaging work environment
- Infrastructure and computational resources for ML model development and research
- Opportunity to work with leading talent on cutting-edge technologies
- Snacks, lunches, dinners, and fun activities
- Competitive compensation package
Tech Stack
Categories
AI ResearchML Engineering
About XPENG
XPENG designs, manufactures, and sells smart electric vehicles for consumers, integrating in-house advanced driver-assistance systems and connected in-car software. Founded in 2014 and headquartered in Guangzhou, it operates plants in Zhaoqing and Guangzhou, maintains a European HQ in Amsterdam, and develops eVTOL aircraft via XPENG AEROHT.