6 months ago
Base Salary
$189k - $290k/yr
Responsibilities
- Design and train Vision-Language-Action solutions for robotaxis.
- Lead end-to-end data strategy, including data mining, auto-labeling, and dataset construction.
- Lead the post-training stack for vision-language models and vision-language-action models, including continual pre-training and supervised fine-tuning.
- Use large-scale data pipelines and machine learning infrastructure to research, prototype, and deploy solutions that improve driving behavior.
- Partner with cross-functional teams to integrate perception signals.
Requirements
- An MS or PhD in Computer Science or a related field is required.
- Experience developing deep learning solutions for vision-language and vision-language-action models is required.
- A track record of post-training large-scale models using continual pre-training, supervised fine-tuning, and reinforcement learning is required.
- Hands-on experience with production machine learning pipelines, dataset creation, training frameworks, and metrics is required.
- Expertise with Python libraries including PyTorch, NumPy, Pandas, and vLLM is required.
- Deep knowledge of cutting-edge computer vision techniques is preferred.
- Publications in top-tier conferences such as CVPR, ICCV, RSS, or ICRA are preferred.
- Experience integrating large language models into various tasks is preferred.
Categories
ML EngineeringRobotics
About Zoox
Zoox builds a fully autonomous, purpose-built electric vehicle and the end-to-end platform to operate it as an urban robotaxi service. The company designs the vehicle, self-driving software, and fleet operations in-house to offer mobility-as-a-service rather than selling cars. Founded in 2014 and headquartered in Foster City, California, Zoox is a subsidiary of Amazon.
