2 months ago
Base Salary
$175k - $296k/yr
Responsibilities
- Research and develop predictive world models that forecast future scene states from multimodal driving and robotics data.
- Develop multi-view future prediction and generation for action-conditioned rollouts and forecasts without explicit action conditioning.
- Create architectures that jointly predict the future and produce trajectories or actions, using predictive pre-training to improve Vision-Language-Action driving performance.
- Extend prediction into multimodal latent spaces covering 3D scene representations, occupancy, and reward signals for simulation, evaluation, and policy training.
- Design unified observation and action representations and embodiment-conditioning mechanisms for transfer across vehicles, robots, and sensor configurations.
- Define evaluation methodology covering representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and closed-loop policy performance.
- Feed evaluated world models back into training as synthetic data and corner-case simulation.
Requirements
- MS or PhD-level education in Engineering or Computer Science focused on deep learning, computer vision, generative models, or a related field, or equivalent experience; recent graduates are welcome.
- Strong applied deep learning experience, including model architecture design, large-scale model training, data curation, and empirical analysis.
- At least 1 year of experience with deep learning frameworks such as PyTorch and hands-on distributed training using FSDP, DeepSpeed, or Megatron-style parallelism.
- Strong Python programming and software design skills.
- Solid understanding of data structures, algorithms, code optimization, and large-scale data processing.
- Ability to design controlled experiments and draw sound conclusions from noisy training signals.
- Preferred experience with generative video or 3D models, including diffusion, flow matching, autoregressive video prediction, NeRF, or Gaussian Splatting.
- Preferred experience with world models or learned simulators for decision-making, model-based reinforcement learning, or Vision-Language-Action models.
- Preferred experience with multimodal foundation models, video tokenizers, VAEs, and pretraining or adapting large pretrained backbones.
- Preferred experience with mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.
Benefits
- Supportive and engaging work environment
- Infrastructure and computational resources to support the work
- Opportunity to work on cutting-edge technologies with leading experts
- Opportunity to contribute to autonomous driving and transportation innovation
- Competitive compensation package with bonus, equity, and benefits
- Snacks, lunches, dinners, and fun activities
Categories
AI ResearchML Engineering
About XPENG
XPENG designs, manufactures, and sells smart electric vehicles for consumers, integrating in-house advanced driver-assistance systems and connected in-car software. Founded in 2014 and headquartered in Guangzhou, it operates plants in Zhaoqing and Guangzhou, maintains a European HQ in Amsterdam, and develops eVTOL aircraft via XPENG AEROHT.