Zoox

AI Inference Engineer - Model Optimization & Deployment

Zoox
Apply
5 months ago
Foster City, CA, USA +2 moreSenior
H1B Sponsor

Base Salary

$242k - $290k/yr

Responsibilities

  • Optimize large-scale LLM and VLM models using quantization, mixed-precision inference, and parameter-efficient fine-tuning workflows.
  • Architect and implement model conversion and compilation pipelines for edge deployment using TensorRT and TensorRT-LLM.
  • Perform parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
  • Write and optimize custom CUDA kernels and TensorRT Plugins to improve memory bandwidth and latency on AI accelerators.
  • Develop production-level, highly concurrent, memory-safe C++ and Python for real-time inference on vehicle system-on-chips.

Requirements

  • Deep expertise in model quantization, including PTQ and QAT, and mixed-precision inference using INT8, FP8, INT4, BF16, and FP16.
  • Experience optimizing LLMs, VLMs, or VLAs with KV-cache optimization, PagedAttention, speculative decoding, FlashAttention, and Linear Attention.
  • Extensive experience with TensorRT and TensorRT-LLM model conversion and compilation pipelines, including parity and latency benchmarking.
  • Proficiency writing and optimizing custom CUDA kernels and TensorRT Plugins for AI accelerators.
  • Production-level C++ 14/17/20 and Python skills, including concurrent, memory-safe, real-time inference code for edge devices.
  • Preferred experience with PyTorch Distributed, Ray, DeepSpeed, and Megatron-LM for distributed training, model parallelism, and GPU-cluster efficiency optimization.
  • Preferred familiarity with autonomous-driving perception stacks, including temporal 3D object detection, BEV, 3D Occupancy Networks, and Vision, LiDAR, and Radar sensor streams.
  • Preferred understanding of VLA models and closed-loop simulation validation for autonomous driving.
Zoox

About Zoox

1,001-5,000 employees

Zoox is transforming mobility-as-a-service by developing a fully autonomous, purpose-built fleet designed for AI to drive and humans to enjoy.

Contact me