Atoms

Staff Cluster Infrastructure Engineer

Atoms
Apply
10 days ago

Base Salary

$224k - $284k/yr

Responsibilities

  • Manage and automate GPU training clusters, including provisioning, bootstrapping, and lifecycle management.
  • Automate bare-metal bring-up so new machines can be added quickly and reliably.
  • Build software abstractions for training and simulation workloads.
  • Improve infrastructure speed, reliability, automation, and uptime at the hardware/software boundary.
  • Diagnose and resolve operational issues in critical systems.
  • Design infrastructure to scale from a small cluster to a larger fleet.

Requirements

  • 6+ years of experience operating GPU compute on Kubernetes or similar orchestration systems.
  • Strong programming and scripting skills in Python, Go, or similar languages.
  • Familiarity with Infrastructure-as-Code tools such as Terraform or CloudFormation.
  • Comfort working with bare-metal Linux environments, GPU hardware, and networking.
  • Experience or demonstrated judgment in scaling critical compute infrastructure, with a focus on automation and reliability.

Benefits

  • Medical, dental, vision, disability, and life insurance.
  • Flexible Spending Account and Health Savings Account options.
  • 401(k), equity eligibility, sick time, unlimited flexible time off, and paid holidays.
  • Paid parental leave and a pre-tax commuter benefit plan.
  • Team lunch in the SoMa office every Tuesday and Thursday.
  • Based in San Francisco and onsite five days per week for office-based teams.
  • Full-time exempt USA employee benefits are subject to change at the company’s discretion.
Atoms

About Atoms

501-1,000 employees
Contact me