10 days ago
Base Salary
$224k - $284k/yr
Responsibilities
- Manage and automate GPU training clusters, including provisioning, bootstrapping, and lifecycle management.
- Automate bare-metal bring-up so new machines can be added quickly and reliably.
- Build software abstractions for training and simulation workloads.
- Improve infrastructure speed, reliability, automation, and uptime at the hardware/software boundary.
- Diagnose and resolve operational issues in critical systems.
- Design infrastructure to scale from a small cluster to a larger fleet.
Requirements
- 6+ years of experience operating GPU compute on Kubernetes or similar orchestration systems.
- Strong programming and scripting skills in Python, Go, or similar languages.
- Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
- Comfort working with bare-metal Linux environments, GPU hardware, and networking.
- Experience or demonstrated judgment in scaling critical compute infrastructure, with a focus on automation and reliability.
Benefits
- Medical, dental, vision, disability, and life insurance.
- Flexible Spending Account and Health Savings Account options.
- 401(k), equity eligibility, sick time, unlimited flexible time off, and paid holidays.
- Paid parental leave and a pre-tax commuter benefit plan.
- Team lunch in the SoMa office every Tuesday and Thursday.
- Based in San Francisco and onsite five days per week for office-based teams.
- Full-time exempt USA employee benefits are subject to change at the company’s discretion.
