2 months ago
Boston, MA, USA +2 moreSenior / Staff+
H1B Sponsor
Base Salary
$226k - $307k/yr
Responsibilities
- Allocate CPU, GPU, and interconnect resources across perception models and inference engines running on the robot.
- Lead initiatives to improve compute utilization through model sharing, model fusion, and scheduling strategies.
- Optimize multimodal sensor-fusion models, LLMs, and VLMs using quantization, pruning, mixed precision, and parameter-efficient fine-tuning.
- Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment.
- Write production-level, low-latency, memory-safe C++ and CUDA code for real-time vehicle inference.
Requirements
- Deep experience optimizing CPU/GPU systems for low latency or high throughput.
- Deep expertise in real-time systems, including latency, memory utilization, and memory bandwidth constraints.
- Deep expertise in PTQ, QAT, and mixed-precision inference using INT8, FP8, FP4, and BF16/FP16.
- Proficiency developing custom ML operators and TensorRT plugins with efficient CUDA kernels for AI accelerators.
- Production-level C++ and Python programming experience, including concurrent, memory-safe, real-time inference code for edge devices.
- Prior high-performance robotics experience in AV, drone, or robot applications is a bonus.
- Familiarity with autonomous-driving perception algorithms, multimodal sensor processing, autonomous-driving foundation models, and edge deployment technologies such as TensorRT-LLM is a bonus.
Categories
About Zoox
Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.