Base Salary
$175k - $296k/yr
Responsibilities
- Research and develop predictive world models that forecast future scene states from multimodal driving and robotics data.
- Develop multi-view future prediction and generation for action-conditioned and non-action-conditioned rollouts.
- Build architectures that jointly predict the future and produce trajectories or actions, using predictive pre-training to improve Vision-Language-Action driving performance.
- Extend prediction into multimodal latent spaces containing 3D scene representations, occupancy, and reward signals for simulation, evaluation, and policy training.
- Design unified observation, action, and embodiment-conditioning mechanisms for transfer across vehicles, robots, and sensor configurations.
- Define evaluation methodologies covering representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon consistency, and closed-loop policy performance.
- Feed evaluated models back into training as synthetic data and corner-case simulation.
Requirements
- MS or PhD-level education in Engineering or Computer Science focused on deep learning, computer vision, generative models, or a related field, or equivalent experience; recent graduates are welcome.
- At least 1 year of experience with applied deep learning, including architecture design, large-scale model training, data curation, and empirical analysis.
- Hands-on experience with PyTorch and distributed training approaches such as FSDP, DeepSpeed, or Megatron-style parallelism.
- Strong Python programming and software design skills.
- Understanding of data structures, algorithms, code optimization, and large-scale data processing.
- Preferred experience with generative video or 3D models, including diffusion, flow matching, autoregressive video prediction, NeRF, or Gaussian Splatting.
- Preferred experience with world models, learned simulators, model-based reinforcement learning, or Vision-Language-Action models.
- Preferred experience with multimodal foundation models, video tokenizers, VAEs, large pretrained backbones, and large-scale training optimization such as mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.
Benefits
- Supportive and engaging work environment
- Infrastructure and computational resources to support the work
- Opportunity to work on cutting-edge technologies and autonomous driving
- Snacks, lunches, dinners, and fun activities
- Bonus, equity, and benefits in addition to base salary
About XPENG
XPENG is a leading Chinese Smart EV company that designs, develops, manufactures, and markets Smart EVs that appeal to the large and growing base of technology-savvy middle-class consumers. Its mission is to drive Smart EV transformation with technology and data, shaping the mobility experience of the future. In order to optimize its customers’ mobility experience, XPeng develops in-house its full-stack advanced driver-assistance system technology and in-car intelligent operating system, as well as core vehicle systems including powertrain and the electrical/electronic architecture. XPeng is headquartered in Guangzhou, China. In 2021, the Company established its European headquarters in Amsterdam, along with other dedicated offices in Copenhagen, Munich, Oslo, and Stockholm.The Company’s Smart EVs are mainly manufactured at its plant in Zhaoqing and Guangzhou,Guangdong province. For more information, please visit https://www.xpeng.com/