5 months ago
Base Salary
$242k - $290k/yr
Responsibilities
- Optimize large-scale LLM and VLM models using quantization, mixed-precision inference, and parameter-efficient fine-tuning workflows.
- Architect and implement model conversion and compilation pipelines for edge deployment using TensorRT and TensorRT-LLM.
- Perform parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
- Write and optimize custom CUDA kernels and TensorRT Plugins to improve memory bandwidth and latency on AI accelerators.
- Develop production-level, highly concurrent, memory-safe C++ and Python for real-time inference on vehicle system-on-chips.
Requirements
- Deep expertise in model quantization, including PTQ and QAT, and mixed-precision inference using INT8, FP8, INT4, BF16, and FP16.
- Experience optimizing LLMs, VLMs, or VLAs with KV-cache optimization, PagedAttention, speculative decoding, FlashAttention, and Linear Attention.
- Extensive experience with TensorRT and TensorRT-LLM model conversion and compilation pipelines, including parity and latency benchmarking.
- Proficiency writing and optimizing custom CUDA kernels and TensorRT Plugins for AI accelerators.
- Production-level C++ 14/17/20 and Python skills, including concurrent, memory-safe, real-time inference code for edge devices.
- Preferred experience with PyTorch Distributed, Ray, DeepSpeed, and Megatron-LM for distributed training, model parallelism, and GPU-cluster efficiency optimization.
- Preferred familiarity with autonomous-driving perception stacks, including temporal 3D object detection, BEV, 3D Occupancy Networks, and Vision, LiDAR, and Radar sensor streams.
- Preferred understanding of VLA models and closed-loop simulation validation for autonomous driving.
Categories
About Zoox
Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.