Deepgram

Applied ML Engineer

Deepgram
Apply
2 months ago
Remote, Worldwide or New York, NY, USASenior / Staff+
H1B Sponsor

Base Salary

$150k - $220k/yr

Responsibilities

  • Own the repeatable pipeline from research checkpoints to deployed, monitored, and scalable production models.
  • Partner with research scientists to turn experimental training and evaluation code into robust, reproducible, and well-tested workflows.
  • Build tooling and abstractions for model training, evaluation, packaging, and deployment.
  • Design automated model release gates for evaluation, regression detection, quality, latency, and throughput.
  • Optimize model inference and serving through batching, memory and latency tuning, profiling, quantization, distillation, compilation, or runtime tuning.
  • Strengthen model build and delivery across GPU compute and cloud environments.
  • Establish end-to-end benchmarking and validation to detect performance and quality regressions.
  • Instrument production model behavior and feed operational results back into research iteration.

Requirements

  • Strong software engineering fundamentals and proficiency in Python, including production-quality, well-tested ML code.
  • Hands-on experience taking machine learning models from research or prototype stage into production at scale.
  • Working understanding of modern deep learning stacks such as PyTorch and the realities of training, evaluating, and serving large models.
  • Experience building ML pipelines and tooling for training orchestration, evaluation, model packaging, deployment, or model CI/CD.
  • Familiarity with serving and inference optimization, including latency, throughput, batching, and resource efficiency.
  • Comfort working across distributed systems and GPU compute in cloud, bare-metal, or hybrid environments.
  • Ability to collaborate with research scientists, scope ambiguous problems, and drive measurable results.
  • Preferred experience with research-to-production handoffs, speech or audio ML, automated model evaluation and release gating, hybrid on-premises and cloud infrastructure, inference optimization, or internal ML platforms and developer tooling.
Deepgram

About Deepgram

51-200 employees

Deepgram is the real-time API platform powering the trillion-dollar Voice AI economy. Backed by a $130M Series C at a $1.3B valuation, Deepgram is trusted by 200,000+ developers and 1,300+ organizations to build Voice AI products, platforms, and autonomous agents with the lowest latency, highest accuracy, and enterprise reliability. Our voice-native foundation models and runtime infrastructure have processed 50,000+ years of audio and over 1 trillion words, making Deepgram the most experienced voice AI platform in the world. Industry-leading models & platform: 👂 Nova-3 — the world’s most accurate real-time speech-to-text model 🔊 Aura-2 — professional, enterprise-grade text-to-speech 💬 Flux — the first Conversational Speech Recognition model designed to handle interruptions 🚀 Voice Agent API — enterprise-ready, real-time conversational AI 🧠 Saga — the Voice OS Beyond core infrastructure, Deepgram is expanding the Voice AI ecosystem through: 💪 Powered by Deepgram, supporting voice products built by leading AI startups and enterprise organizations 🌉 A new Voice AI Collaboration Hub in San Francisco for builders, partners, and the voice community 🍔 The acquisition of OfOne, delivering real-time Voice AI for restaurants and drive-thru operations with 95%+ containment 📃 A growing patent portfolio in Voice AI Much like APIs powered the payments and cloud economies, Deepgram is building the foundation for a trillion-dollar B2B Voice AI economy—centered on the most natural human interface: voice.