6 months ago
Remote, United States or New York, NY, USAMid Level
H1B sponsor
Base Salary
$160k - $220k/yr
Responsibilities
- Design and build CI/CD pipelines for ML model development, validation, and deployment.
- Architect and maintain pipelines that move models from research through staging to production.
- Build A/B testing infrastructure for controlled model rollouts and real-world performance measurement.
- Implement monitoring for production model accuracy, latency, drift, and regressions.
- Develop automated retraining pipelines triggered by data changes, performance degradation, or schedules.
- Create production-mirroring build and test environments for high-fidelity researcher feedback.
- Establish model versioning, artifact management, and rollback capabilities.
- Collaborate with research engineers to define and enforce model quality gates.
- Build observability dashboards for real-time model health across environments.
- Optimize model-serving infrastructure for latency, throughput, and cost efficiency.
Requirements
- At least 4 years of experience in MLOps, DevOps, or infrastructure engineering focused on ML systems.
- Strong proficiency in Python and experience building automation and tooling for ML workflows.
- Deep experience with CI/CD systems and pipelines for software and model delivery.
- Hands-on experience with Docker and Kubernetes for containerized workloads.
- Practical experience deploying and serving ML models in production.
- Familiarity with model evaluation, validation, and quality assurance processes.
- Understanding of monitoring and observability principles for ML systems.
- Experience with model serving frameworks such as NVIDIA Triton Inference Server, TensorRT, or ONNX Runtime is preferred.
- Background in speech, audio, or real-time media ML systems is preferred.
- Experience with Terraform or Pulumi is preferred.
- Experience with Prometheus, Grafana, Datadog, or similar monitoring and observability stacks is preferred.
- Familiarity with GPU-accelerated inference optimization, profiling, feature stores, data versioning, ML metadata management, canary deployments, or progressive delivery is preferred.
Benefits
- Medical, dental, and vision benefits
- Annual wellness stipend and mental health support
- Life, short-term disability, and long-term disability insurance plans
- Unlimited paid time off, generous paid parental leave, flexible schedule, and 12 paid US company holidays
- Quarterly personal productivity stipend and one-time home office upgrade stipend
- 401(k) plan with company match and tax savings programs
- Learning and education stipend, conference participation, employee resource groups, and AI enablement workshops
- Benefits for international employees are administered locally through an Employer of Record and vary by region
Tech Stack
Categories
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
