3 months ago
Remote, United States or New York, NY, USASenior / Staff+
H1B sponsor
Base Salary
$150k - $220k/yr
Responsibilities
- Own the repeatable pipeline from research checkpoints to deployed, monitored, and scalable production models.
- Partner with research scientists to turn experimental training and evaluation code into robust, reproducible, and well-tested workflows.
- Build tooling and abstractions for model training, evaluation, packaging, and deployment.
- Design automated model release gates for evaluation, regression detection, quality, latency, and throughput.
- Optimize model inference and serving through batching, memory and latency tuning, profiling, quantization, distillation, compilation, or runtime tuning.
- Strengthen model build and delivery across GPU compute and cloud environments.
- Establish end-to-end benchmarking and validation to detect performance and quality regressions.
- Instrument production model behavior and feed operational results back into research iteration.
Requirements
- Strong software engineering fundamentals and proficiency in Python, including production-quality, well-tested ML code.
- Hands-on experience taking machine learning models from research or prototype stage into production at scale.
- Working understanding of modern deep learning stacks such as PyTorch and the realities of training, evaluating, and serving large models.
- Experience building ML pipelines and tooling for training orchestration, evaluation, model packaging, deployment, or model CI/CD.
- Familiarity with serving and inference optimization, including latency, throughput, batching, and resource efficiency.
- Comfort working across distributed systems and GPU compute in cloud, bare-metal, or hybrid environments.
- Ability to collaborate with research scientists, scope ambiguous problems, and drive measurable results.
- Preferred experience with research-to-production handoffs, speech or audio ML, automated model evaluation and release gating, hybrid on-premises and cloud infrastructure, inference optimization, or internal ML platforms and developer tooling.
Categories
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
