2 months ago
Base Salary
$180k - $240k/yr
Responsibilities
- Define evaluation methodologies for speech-to-text, text-to-speech, LLM, RAG, agent, and multimodal systems.
- Build and maintain automated batch and streaming evaluation pipelines for quality, hallucination, latency, and time-to-first-byte detection.
- Develop scalable evaluation infrastructure including harnesses, orchestration, and result-aggregation pipelines running against production models and GPU clusters.
- Convert Research benchmarks and model metrics into automated pass/fail quality gates.
- Build and operate canaries and continuous-monitoring systems to detect production quality regressions.
- Partner with DevOps and Infrastructure to create ephemeral test environments and results-aggregation infrastructure.
- Collaborate with Research, model training, inference, and product teams to provide evaluation signals for release and optimization decisions.
- Integrate evaluation and quality gates into CI/CD.
- Contribute through code reviews, technical design discussions, and engineering and QA practices.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science, AI, Applied Math, or a related field, or equivalent experience.
- At least 5 years of professional software or QA engineering experience, including shipping test infrastructure or evaluation systems.
- Backend or scripting experience with Python, Rust, Go, or a similar language.
- Experience designing automated test pipelines, evaluation frameworks, or data-processing systems.
- Strong analytical skills and ability to reason about metrics, thresholds, statistical variation, and regression detection.
- Ability to lead ambiguous technical work and communicate across research, engineering, and product teams.
- Preferred experience evaluating LLMs, RAG pipelines, agents, or multimodal models.
- Preferred experience with React Native or other cross-platform mobile frameworks.
- Preferred experience building evaluation frameworks, benchmarks, or ML infrastructure for internal teams or external users.
- Preferred experience with voice, audio, speech recognition, real-time systems, WER, MOS, or latency/TTFB.
- Open-source contribution, review, maintenance, or community engagement experience is a plus.
- Experience bridging evaluation, training, inference, or agent-framework teams is a plus.
- Familiarity with cloud infrastructure, containerized or ephemeral environments, Grafana, canaries, and anomaly detection is a plus.
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
