Baseten

Global Capacity Manager

Baseten
Apply
5 hours ago

Base Salary

$170k - $230k/yr

Responsibilities

  • Lead specialized GPU pods through acquisition, orchestration, maintenance, and the full asset lifecycle.
  • Architect infrastructure readiness and deployment strategies for GPU clusters, including NVIDIA Blackwell B200 systems.
  • Build multi-cloud capacity management systems and execute workload migrations and deployment drains across regions.
  • Design and implement capacity management systems capable of supporting a 10x increase in GPU volume.
  • Develop automated Go-based operators to identify, cordon, and repair unhealthy GPU nodes.
  • Build ROI and financial models for GPU spending and capacity-versus-reliability trade-offs.
  • Partner with SRE, infrastructure, and FDE teams to complete infrastructure changes and operational tasks.
  • Lead capacity-crunch incident response by re-coordinating workloads during outages.

Requirements

  • Bachelor’s, Master’s, or Ph.D. degree in Computer Science, Engineering, Mathematics, or a related field.
  • 5+ years of professional experience in a high-growth environment, preferably at a hyperscaler or specialized GPU provider.
  • Deep hands-on Kubernetes expertise, including taints, cordons, node draining, and custom operators.
  • Production-level experience with Go or Python.
  • Strong financial literacy and ability to model trade-offs between capacity reliability and cost.
  • High tenacity and a collaborative mindset.

Benefits

  • Competitive compensation including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employees and dependents.
  • Flexible PTO and company-wide Winter Break from Christmas Eve through New Year’s Day.
  • Paid parental leave.
  • Fertility and family-building stipend through Carrot.
  • Company-facilitated 401(k).
  • Exposure to a variety of ML startups and related learning and networking opportunities.
Baseten

About Baseten

201-500 employees

Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.

Contact me