4 months ago
Remote, Americas +2 moreStaff+
Base Salary
$195k - $286k/yr
Responsibilities
- Define and drive technical strategy for model distillation and compression across Waabi’s AI stack.
- Design, implement, and scale distillation and model-efficiency pipelines for generative and other large-scale neural network models.
- Develop quantization-aware and post-training quantization, knowledge distillation, pruning, sparsification, low-rank factorization, efficient architecture, and inference-time efficiency techniques.
- Integrate compressed models into production onboard autonomy and high-throughput simulation pipelines while meeting latency, memory, and throughput targets.
- Define benchmarks and evaluation frameworks for efficiency versus quality trade-offs across models and hardware targets.
- Partner with ML Platform, Infrastructure, Onboard Autonomy, and Simulation teams.
- Mentor researchers and engineers and set a high technical bar for model efficiency work.
- Champion model-compression best practices through design reviews, documentation, and technical talks.
- Contribute to scientific publications and open-source projects.
Requirements
- Extensive hands-on experience designing and implementing distillation, quantization, pruning, and model-compression techniques for large-scale neural networks with demonstrated production impact.
- Bachelor’s or Master’s degree in Machine Learning, Computer Vision, Robotics, or a related field, or equivalent industry experience.
- Expert Python and PyTorch or JAX skills and experience with large-scale distributed training.
- Proven track record of setting technical direction and driving projects from conception to production.
- Experience collaborating with infrastructure, platform, and autonomy teams to deploy compressed models under engineering constraints.
- Ability to communicate complex technical trade-offs clearly and drive alignment across research and engineering teams.
- Preferred experience with hardware-aware optimization, including TensorRT, ONNX, custom CUDA kernels, and hardware-specific quantization.
- Preferred publications at top-tier ML or computer vision venues such as NeurIPS, ICML, CVPR, ICLR, or ECCV.
- Preferred experience distilling diffusion models, LLMs, VLMs, or video models.
- Preferred background in autonomous vehicles or robotics.
Benefits
- Competitive compensation and equity awards.
- Medical, dental, and vision coverage for full-time employees.
- Unlimited vacation.
- Flexible hours and work-from-home support.
- Daily drinks, snacks, and catered meals when in the office.
- Regular on-site, off-site, and virtual team-building activities and social events.
- US-based role with locations across Waabi’s offices; work-from-home support is available.
