10 days ago
San Jose, CA, USAMid Level
Base Salary
$120k - $300k/yr
Responsibilities
- Design response, feature, and self-distillation strategies to compress teacher models into deployable student models.
- Apply post-training quantization, quantization-aware training, pruning, and architecture search to meet product size, latency, and power budgets.
- Build a reusable distillation and compression toolchain for the audio ML team.
- Partner with audio ML and runtime teams on training pipelines and deployment targets.
- Define and track accuracy-retention and resource KPIs through the release cycle.
- Profile compressed models on target hardware and iterate with DSP and runtime engineers on bottlenecks.
Requirements
- 3+ years of professional experience in model compression, distillation, quantization, or efficient deep learning.
- Strong fluency in PyTorch or TensorFlow and modern compression libraries.
- Hands-on experience converting models from full precision to fixed-point or int8 with controlled accuracy loss.
- Ability to reason about compute, memory bandwidth, and power constraints close to hardware.
- Track record of producing models shipped to constrained devices.
- Strong foundation in audio or sequence model architectures, including CNNs, transformers, RNN-T, or conformers.
Tech Stack
PyTorchTensorFlow
