10 days ago
San Jose, CA, USASenior
Base Salary
$180k - $450k/yr
Responsibilities
- Design and maintain Infrastructure as Code practices for repeatable and scalable cluster provisioning.
- Enhance deployment pipelines for secure, reliable, and low-latency model service delivery.
- Own training infrastructure operating at the scale of 10,000+ GPUs, including scheduling, fault tolerance, and network optimization.
- Partner with ML researchers and engineers to address compute bottlenecks through infrastructure improvements.
- Monitor system health, define SLOs, and lead incident response for critical workloads.
- Drive capacity planning, cost efficiency, and GPU hardware lifecycle management.
- Build internal tooling and platform abstractions that improve developer experience for compute users.
Requirements
- 5+ years of experience in infrastructure, systems, or platform engineering, including at least 2 years in ML or HPC environments.
- Experience managing GPU clusters or large-scale distributed compute infrastructure.
- Strong proficiency in at least one systems or infrastructure programming language.
- Deep understanding of networking fundamentals relevant to high-throughput training workloads, with RDMA, InfiniBand, or RoCE experience advantageous.
- Experience with container orchestration, job scheduling, and multi-tenant resource management.
- Track record of owning production systems with high reliability requirements.
- Strong debugging and observability skills across the infrastructure stack.
- Bonus experience includes operating large GPU-aware Kubernetes clusters, using Pulumi or similar IaC tooling, developing systems-level tooling in Rust or Go, and familiarity with PyTorch and Ray.
Benefits
- Full-time position with a US base salary range of $180,000-$450,000 annually.
- Additional compensation components and benefits may be included depending on the specific role.
