Figure AI

Helix AI Engineer, Video Pretraining

Figure AI
Apply
1 month ago
San Jose, CA, USASenior

Base Salary

$200k - $400k/yr

Responsibilities

  • Design and train large-scale video foundation models on internet-scale and robot-collected datasets.
  • Develop pretraining strategies that capture temporal dynamics, motion, and object interaction from raw video.
  • Build transferable representations for perception, tracking, prediction, and control.
  • Explore transformer-based and diffusion-based architectures for video understanding and generation.
  • Implement efficient video data pipelines and distributed training strategies.
  • Optimize model performance across compute, memory, and training-efficiency constraints.
  • Collaborate with generative modeling, agent, and robot learning teams to integrate pretrained models into the autonomy stack.
  • Design evaluation frameworks and benchmarks for temporal understanding, prediction quality, and generalization.

Requirements

  • Experience training large-scale models on video data or other high-dimensional sequential modalities.
  • Strong understanding of modern deep learning architectures for video, vision, or multimodal systems.
  • Experience with large-scale pretraining, dataset curation, training dynamics, and scaling laws.
  • Proficiency in Python and deep learning frameworks such as PyTorch.
  • Experience with distributed training systems and large GPU clusters.
  • Strong experimental rigor and ability to iterate quickly on model design and training strategies.
  • Solid software engineering skills and ability to build scalable, reliable systems.
  • Ability to work independently and drive ambiguous, high-impact research directions.
  • Preferred qualifications include experience with frontier video or multimodal foundation models, video diffusion, autoregressive video modeling, world models, large-scale video dataset construction and filtering, robotics or embodied AI, egocentric video, and publications in machine learning, computer vision, or multimodal AI.
  • Experience at leading AI labs such as OpenAI, Google DeepMind, Google, ByteDance, Midjourney, or Adobe is a bonus.

Benefits

  • Full-time position requiring five days per week of in-office collaboration in San Jose, California.
  • Total compensation may include additional components and benefits depending on the specific role.

Tech Stack

Categories

AI ResearchML Engineering
Contact me