XPENG

Senior Machine Learning Engineer – Predictive World Model

XPENG
Apply
2 months ago
Santa Clara, CA, USASenior
H1B sponsor

Base Salary

$175k - $296k/yr

Responsibilities

  • Research and develop predictive world models that forecast future scene states from multimodal driving and robotics data.
  • Develop multi-view future prediction and generation for action-conditioned rollouts and forecasts without explicit action conditioning.
  • Create architectures that jointly predict the future and produce trajectories or actions, using predictive pre-training to improve Vision-Language-Action driving performance.
  • Extend prediction into multimodal latent spaces covering 3D scene representations, occupancy, and reward signals for simulation, evaluation, and policy training.
  • Design unified observation and action representations and embodiment-conditioning mechanisms for transfer across vehicles, robots, and sensor configurations.
  • Define evaluation methodology covering representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and closed-loop policy performance.
  • Feed evaluated world models back into training as synthetic data and corner-case simulation.

Requirements

  • MS or PhD-level education in Engineering or Computer Science focused on deep learning, computer vision, generative models, or a related field, or equivalent experience; recent graduates are welcome.
  • Strong applied deep learning experience, including model architecture design, large-scale model training, data curation, and empirical analysis.
  • At least 1 year of experience with deep learning frameworks such as PyTorch and hands-on distributed training using FSDP, DeepSpeed, or Megatron-style parallelism.
  • Strong Python programming and software design skills.
  • Solid understanding of data structures, algorithms, code optimization, and large-scale data processing.
  • Ability to design controlled experiments and draw sound conclusions from noisy training signals.
  • Preferred experience with generative video or 3D models, including diffusion, flow matching, autoregressive video prediction, NeRF, or Gaussian Splatting.
  • Preferred experience with world models or learned simulators for decision-making, model-based reinforcement learning, or Vision-Language-Action models.
  • Preferred experience with multimodal foundation models, video tokenizers, VAEs, and pretraining or adapting large pretrained backbones.
  • Preferred experience with mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.

Benefits

  • Supportive and engaging work environment
  • Infrastructure and computational resources to support the work
  • Opportunity to work on cutting-edge technologies with leading experts
  • Opportunity to contribute to autonomous driving and transportation innovation
  • Competitive compensation package with bonus, equity, and benefits
  • Snacks, lunches, dinners, and fun activities

Tech Stack

Categories

AI ResearchML Engineering
XPENG

About XPENG

1,001-5,000 employees

XPENG designs, manufactures, and sells smart electric vehicles for consumers, integrating in-house advanced driver-assistance systems and connected in-car software. Founded in 2014 and headquartered in Guangzhou, it operates plants in Zhaoqing and Guangzhou, maintains a European HQ in Amsterdam, and develops eVTOL aircraft via XPENG AEROHT.

Contact me