
Staff Cluster Infrastructure Engineer
CloudKitchens12 days ago
Base Salary
$224k - $284k/yr
Responsibilities
- Manage and automate GPU training clusters, including provisioning, bootstrapping, and lifecycle management.
- Automate bare-metal bring-up so new machines become operational quickly and reliably.
- Build software abstractions that provide a unified interface to training and simulation workloads.
- Improve automation, performance, reliability, and uptime at the hardware/software boundary.
- Diagnose and resolve operational issues in critical systems.
- Design infrastructure to scale from a smaller cluster to a larger fleet of machines.
Requirements
- 6+ years of experience operating GPU compute on Kubernetes or similar orchestration systems.
- Strong programming and scripting skills in Python, Go, or similar languages.
- Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
- Comfort working with bare-metal Linux environments, GPU hardware, and networking.
- Demonstrated focus on automation, reliability, and operating critical systems.
Benefits
- Medical, dental, vision, disability, and life insurance.
- Flexible Spending Account and Health Savings Account options.
- 401(k), equity eligibility, sick time, unlimited flexible time off, and paid holidays.
- Paid parental leave and a pre-tax commuter benefit plan.
- Team lunch in the SoMa office every Tuesday and Thursday.
- Based in San Francisco and onsite five days per week for office-based teams.
About CloudKitchens
We provide kitchen infrastructure and software that empower food & beverage operators to expand their operations with minimal upfront capital and time.