5 months ago
Base Salary
$189k - $290k/yr
Responsibilities
- Design and train Vision-Language-Action solutions for robotaxis.
- Lead end-to-end data strategy, including data mining, auto-labeling, and dataset construction.
- Lead the post-training stack for vision-language models and vision-language-action models, including continual pre-training and supervised fine-tuning.
- Use large-scale data pipelines and machine learning infrastructure to research, prototype, and deploy solutions that improve driving behavior.
- Partner with cross-functional teams to integrate perception signals.
Requirements
- An MS or PhD in Computer Science or a related field is required.
- Experience developing deep learning solutions for vision-language and vision-language-action models is required.
- A track record of post-training large-scale models using continual pre-training, supervised fine-tuning, and reinforcement learning is required.
- Hands-on experience with production machine learning pipelines, dataset creation, training frameworks, and metrics is required.
- Expertise with Python libraries including PyTorch, NumPy, Pandas, and vLLM is required.
- Deep knowledge of cutting-edge computer vision techniques is preferred.
- Publications in top-tier conferences such as CVPR, ICCV, RSS, or ICRA are preferred.
- Experience integrating large language models into various tasks is preferred.
Categories
ML EngineeringRobotics
About Zoox
Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.