27 days ago
Base Salary
$242k - $290k/yr
Responsibilities
- Design and implement multimodal Lidar, Camera, and Radar sensor-fusion architectures for 3D occupancy, semantic segmentation, and flow prediction.
- Develop vision-first fusion strategies to improve geometric understanding and reduce reliance on sparse sensor modalities.
- Engineer temporal processing modules for stable and consistent predictions over time.
- Optimize model architectures for real-time on-vehicle inference under strict latency constraints.
- Collaborate with Tracking, Prediction, and Planner teams to refine contours, free-space estimates, and other geometric outputs for complex maneuvers.
Requirements
- MS or PhD in Computer Science, Robotics, Machine Learning, or a related field.
- 6+ years of industry experience.
- Deep expertise in 3D computer vision and deep learning, particularly voxel-based or BEV architectures.
- Strong proficiency in Python and PyTorch, plus some C++ experience for model integration.
- Experience with Lidar, Camera, and Radar sensor fusion and temporal data sequences.
- Experience with occupancy networks, implicit representations such as NeRF or Gaussian Splats, or scene-flow estimation.
- Bonus: experience optimizing models with TensorRT and CUDA for low-latency inference.
- Bonus: familiarity with sparse convolutions, query-based architectures, Vision Language Models, multimodal 3D foundation models, World Models, or VLA.
Categories
ML EngineeringRobotics
About Zoox
Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.