6 months ago
Base Salary
$160k - $220k/yr
Responsibilities
- Define and drive end-to-end infrastructure architecture for production inference and research training AI/ML workloads.
- Design multi-cloud and hybrid infrastructure strategies balancing performance, reliability, cost, and vendor flexibility.
- Architect compute orchestration systems for GPU and CPU workloads across heterogeneous infrastructure.
- Design storage architectures for high-throughput training data pipelines and low-latency model serving.
- Lead capacity planning and growth modeling across infrastructure dimensions.
- Drive infrastructure cost optimization and FinOps practices.
- Design burstable, elastic training infrastructure that scales with demand while minimizing idle cost.
- Architect research compute infrastructure for ML teams while maintaining operational efficiency.
- Establish architectural standards, design review processes, and technical documentation practices.
- Collaborate with engineering leadership to align infrastructure strategy with product and business objectives.
- Evaluate emerging hardware, cloud services, and infrastructure technologies.
Requirements
- 7+ years of experience in infrastructure engineering, systems architecture, or a senior technical role focused on large-scale infrastructure.
- Proven experience designing multi-cloud architectures spanning AWS and at least one other major cloud provider or on-premises environment.
- Deep expertise in block, object, and file storage, including performance tuning for large-scale data workloads.
- Strong experience with compute orchestration using Kubernetes and workload scheduling across heterogeneous infrastructure.
- Hands-on experience with GPU infrastructure, including procurement considerations, cluster design, driver management, and runtime management.
- Track record of capacity planning and infrastructure scaling for high-growth environments.
- Ability to communicate complex architectural decisions clearly to technical and non-technical stakeholders.
- Strong understanding of networking fundamentals relevant to infrastructure architecture.
- Preferred: experience architecting ML training infrastructure, including distributed training, large dataset management, and experiment infrastructure.
- Preferred: background in cost optimization and FinOps for large-scale cloud and bare metal infrastructure.
- Preferred: experience operating bare metal infrastructure in colocation facilities.
- Preferred: expertise in network architecture, high-bandwidth GPU interconnects, and global traffic routing.
- Preferred: experience with infrastructure modeling and simulation for capacity planning.
- Preferred: familiarity with Slurm, Ray, or other HPC/ML job scheduling systems.
- Preferred: understanding of power, cooling, and physical infrastructure considerations for GPU-dense deployments.
Benefits
- Medical, dental, and vision benefits; annual wellness stipend; mental health support; life, STD, and LTD income insurance plans.
- Unlimited PTO, generous paid parental leave, flexible schedule, and 12 paid US company holidays.
- Quarterly personal productivity stipend, one-time home office upgrade stipend, 401(k) with company match, and tax savings programs.
- Learning and education stipend, participation in talks and conferences, employee resource groups, and AI enablement workshops.
- Benefits for international employees are administered locally through an Employer of Record and vary by region.
Tech Stack
Categories
About Deepgram
Deepgram builds Voice AI infrastructure for developers, offering speech-to-text, text-to-speech, and voice agent tooling through real-time APIs and deployable on-prem software. Privately held and founded in 2015, the company is headquartered in San Francisco. Customers such as Twilio, Cloudflare, and Jack in the Box use its models to power production voice features; products include Nova-3 STT, Aura TTS, and a Voice Agent API.
