1 month ago
San Jose, CA, USASenior
Base Salary
$200k - $400k/yr
Responsibilities
- Design, train, and deploy large-scale diffusion and other generative models for vision, video, and multimodal data.
- Develop models for robot perception, world modeling, and prediction from raw sensory inputs.
- Build systems for synthetic data creation, augmentation, and dataset scaling for robot learning.
- Research and implement state-of-the-art diffusion, generative modeling, and multimodal foundation model techniques.
- Optimize distributed training pipelines and integrate generative models with data, infrastructure, and agent teams.
- Evaluate model quality, robustness, and generalization and contribute to scalable experimentation frameworks.
Requirements
- Experience training and deploying diffusion, autoregressive, or related generative models at scale.
- Strong understanding of modern deep learning for vision and/or multimodal systems.
- Proficiency in Python and deep learning frameworks such as PyTorch.
- Experience with large-scale datasets and distributed training systems.
- Strong experimental rigor, rapid iteration skills, and solid software engineering ability.
- Ability to build reliable, maintainable systems and independently own ambiguous, high-impact technical problems.
- Preferred qualifications include diffusion-based image or video generation, multimodal foundation models, synthetic data or robotics simulation, large-scale multi-node or GPU-cluster training, 3D, video prediction, world models, robotics, embodied AI, real-world machine learning systems, or relevant publications.
Benefits
- Full-time position requiring five days per week of in-office collaboration in San Jose, California.
- Additional compensation components and benefits may be included depending on the specific role.
