10 days ago
San Jose, CA, USAStaff+
Base Salary
$120k - $300k/yr
Responsibilities
- Implement and train audio models for wake-word detection, voice activity detection, source separation, speech enhancement, and related applications.
- Move audio models from research prototypes to on-device deployment within latency, memory, and power budgets.
- Build and maintain training data pipelines, evaluation harnesses, and retraining processes across model families.
- Partner with DSP and firmware engineers to integrate models into the Hark Audio Engine and DSP runtime.
- Collaborate with hardware and acoustics teams to characterize operating signal conditions.
- Profile and optimize models on DSP, NPU, and CPU target platforms and define product accuracy and resource budgets.
Requirements
- At least 3 years of professional experience building and shipping audio or speech ML models.
- Strong fluency with PyTorch or TensorFlow and modern audio deep learning toolchains.
- Hands-on experience deploying models to embedded targets such as DSP, NPU, or mobile NPU and CPU.
- Experience across the full ML lifecycle, including data, training, evaluation, deployment, and monitoring.
- Solid foundation in audio signal processing and its intersection with ML pipelines.
- Experience collaborating with DSP, firmware, and hardware engineers on resource-constrained systems.
- Preferred qualifications include experience shipping voice-first or far-field audio products, on-device wake-word systems, ASR front-ends, or production-scale speech enhancement.
- Preferred qualifications include familiarity with quantization, pruning, distillation, Qualcomm AI stacks or similar providers, and open-source audio ML contributions.
Tech Stack
PyTorchTensorFlow
