4 hours ago
Remote, WorldwideSenior
Base Salary
$96k - $96k/yr
Responsibilities
- Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.
- Measure and reduce latency toward a first-audio target below 800 milliseconds on real calls.
- Handle interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.
- Build an evaluation harness using recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.
- Compare voice providers and models through evidence-based testing and apply the results to the product.
- Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.
- Work directly with founders and make technical decisions in a fast-moving team.
Requirements
- At least 5 years of production software development experience, including 2 or more years shipping voice, speech, or real-time audio systems.
- Experience building and shipping end-to-end real-time voice pipelines covering streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.
- Strong Python or TypeScript skills with comfort working in both languages.
- Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.
- Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.
- Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results for product decisions.
- Clear written English for asynchronous communication.
- Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.
Benefits
- $96,000 USD annual compensation regardless of location.
- Fully remote role open anywhere in the world.
- Core team overlap is 13:00 to 17:00 UTC.
