7 months ago
Base Salary
$150k - $220k/yr
Responsibilities
- Architect and maintain Kubernetes-based computing platforms across AWS and on-premise environments.
- Develop and manage reproducible, versioned, and automated infrastructure using Terraform and Infrastructure-as-Code principles.
- Design and optimize AI/ML job scheduling and orchestration by integrating Slurm with Kubernetes clusters.
- Provision, manage, and maintain bare metal server infrastructure for high-performance GPU computing.
- Implement networking and storage solutions, including CNI, service mesh, CSI, and S3, for hybrid workloads.
- Build observability capabilities covering monitoring, logging, and tracing.
- Create automation for operational tasks, incident response, and performance tuning.
- Collaborate with AI researchers and ML engineers to build infrastructure tools and development workflows.
- Automate the lifecycle of single-tenant managed deployments.
Requirements
- At least 5 years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
- Hands-on experience building and managing production infrastructure with Terraform.
- Expert-level knowledge of Kubernetes architecture and operations in large-scale environments.
- Experience with HPC job schedulers, specifically Slurm, for GPU-intensive AI workloads.
- Experience managing bare metal infrastructure, including provisioning with tools such as PXE boot and MAAS, configuration, and lifecycle management.
- Strong scripting and automation skills using technologies such as Python, Go, or Bash.
- Experience with CI/CD systems such as GitLab CI, Jenkins, or ArgoCD is preferred.
- Experience building developer tooling is preferred.
- Familiarity with FinOps principles and cloud cost optimization is preferred.
- Knowledge of Kubernetes networking solutions such as Calico and Cilium is preferred.
- Knowledge of Kubernetes storage solutions such as Ceph and Rook is preferred.
- Experience in multi-region or hybrid cloud environments is preferred.
Benefits
- Medical, dental, and vision benefits
- Annual wellness stipend
- Mental health support
- Life, short-term disability, and long-term disability income insurance plans
- Unlimited paid time off
- Paid parental leave
- Flexible schedule
- 12 paid U.S. company holidays
- Quarterly personal productivity stipend
- One-time home office upgrade stipend
- 401(k) plan with company match
- Tax savings programs
- Learning and education stipend
- Participation in talks and conferences
- Employee Resource Groups
- AI enablement workshops and sessions
- Benefits are administered locally through an Employer of Record model in many countries and vary by region
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
