10 days ago
San Jose, CA, USASenior
Base Salary
$170k - $400k/yr
Responsibilities
- Own client audio quality, including echo, self-interruption, dropouts, and clipping.
- Build and tune browser audio pipelines with Web Audio API, AudioWorklet, and getUserMedia constraints.
- Work across the WebRTC audio path, including acoustic echo cancellation, noise suppression, and voice activity detection.
- Ship production DSP to the client using C++ or Rust compiled to WebAssembly, as well as TypeScript in the audio pipeline.
- Tune endpointing, interruption handling, and turn-taking for natural voice-agent conversations.
- Reduce conversational latency and audio artifacts across the streaming pipeline.
- Contribute to the React and TypeScript client where audio meets the user interface.
- Manage features end-to-end from prototyping through production.
- Collaborate with designers, platform engineers, and the speech team.
Requirements
- At least 5 years of software engineering experience.
- Experience shipping real-time audio into a product used by real users.
- Hands-on experience with WebRTC, acoustic echo cancellation, noise suppression, and voice activity detection.
- Strong DSP fundamentals, including adaptive filtering, STFT, resampling, and gain control.
- Production experience with C/C++ or Rust DSP and shipping it to browsers through WebAssembly.
- Working knowledge of the browser audio stack, including Web Audio API, AudioWorklet, and MediaStream constraints.
- Comfort working with latency, buffering, and sample rates in streaming audio pipelines.
- Ability to own features end-to-end and work in a shared production codebase.
- Preferred experience at a voice, speech, or video-conferencing company.
- Preferred experience with audio ML, including noise suppression, VAD, or source separation, and on-device inference.
- Preferred familiarity with WebRTC internals, including the Audio Processing Module, AEC3, and Opus.
- Preferred familiarity with voice-agent frameworks such as LiveKit and Pipecat.
- Preferred TypeScript and React experience and experience with target-speaker isolation, diarization, barge-in, or conversational turn-detection systems.
