5 hours ago
Base Salary
$170k - $230k/yr
Responsibilities
- Lead specialized GPU pods through acquisition, orchestration, maintenance, and the full asset lifecycle.
- Architect infrastructure readiness and deployment strategies for GPU clusters, including NVIDIA Blackwell B200 systems.
- Build multi-cloud capacity management systems and execute workload migrations and deployment drains across regions.
- Design and implement capacity management systems capable of supporting a 10x increase in GPU volume.
- Develop automated Go-based operators to identify, cordon, and repair unhealthy GPU nodes.
- Build ROI and financial models for GPU spending and capacity-versus-reliability trade-offs.
- Partner with SRE, infrastructure, and FDE teams to complete infrastructure changes and operational tasks.
- Lead capacity-crunch incident response by re-coordinating workloads during outages.
Requirements
- Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
- 5+ years of professional experience in a high-growth environment, preferably at a hyperscaler or specialized GPU provider.
- Deep hands-on Kubernetes expertise, including taints, cordons, node draining, and custom operators.
- Production-level experience with Go or Python.
- Strong financial literacy and ability to model trade-offs between capacity reliability and cost.
- High tenacity and a collaborative mindset.
Benefits
- Competitive compensation including meaningful equity.
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible PTO and company-wide Winter Break from Christmas Eve through New Year’s Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups and related learning and networking opportunities.
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
