1 day ago
Base Salary
$184k - $357k/yr
Responsibilities
- Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-premises environments.
- Develop tools and services for cluster provisioning, validation, upgrades, monitoring, repair, and lifecycle operations.
- Improve Day 0, Day 1, and Day 2 workflows for cluster bringup, handoff, and production operations.
- Reduce manual production work through APIs, automation, and agent-assisted workflows.
- Participate in on-call, incident response, production debugging, and durable follow-up work.
- Collaborate with platform, storage, networking, security, and workload teams to make infrastructure production-ready.
Requirements
- 8+ years of experience building or operating production infrastructure.
- Strong programming skills in Python, Go, or similar languages.
- Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
- Ability to troubleshoot distributed systems in production.
- Clear communication and ability to work across teams.
- BS/MS in Computer Science or equivalent experience.
- Preferred experience includes GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, fleet automation, SLOs, on-call, incident response, observability, reliability practices, BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure.
Benefits
- Eligible for equity and benefits.
- Applications will be accepted at least until September 7, 2026.
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
