7 hours ago
San Francisco, CA, USAMid Level
Base Salary
$150k - $200k/yr
Responsibilities
- Build and iterate on a cloud-based ASR pipeline from audio capture through post-processing at production scale.
- Own ASR quality and reliability while improving latency, small-word accuracy, and voice-print reliability.
- Work across data preparation, model training and fine-tuning, evaluation, and deployment to ship pipeline improvements.
- Optimize latency-sensitive and streaming audio/ASR pipelines based on real user feedback.
- Collaborate with product engineering, China-based R&D, hardware, and supply-chain teams across time zones.
- Turn loosely defined requirements into concrete, production-ready improvements with minimal team support.
Requirements
- At least 3 years of experience building and tuning production transcription or ASR pipelines end to end, primarily in cloud environments.
- Demonstrated ownership of production ASR systems across data preparation, model training and fine-tuning, evaluation, and deployment.
- Experience with latency-sensitive or streaming audio/ASR pipelines and production lifecycle improvements.
- Early-stage or founding engineering experience with the ability to work independently and ship without fully specified requirements.
- On-device or embedded ML experience with Core ML, TensorFlow Lite, or similar frameworks.
- Prior experience building wearable, hardware, or robotics products.
- Background in AI-native consumer applications focused on transcription or audio.
- Experience building agent or LLM-based product features involving tool use, memory, or retrieval systems.
- Ability to collaborate asynchronously across time zones.
Benefits
- Annual salary range of $150,000 to $200,000 USD.
- Hybrid work arrangement requiring three days per week in the San Francisco Bay Area office.
