24 hours ago
Base Salary
$184k - $357k/yr
Responsibilities
- Build and operate automation for large-scale GPU clusters across cloud partner and on-premises environments.
- Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations.
- Improve Day 0, Day 1, and Day 2 cluster bringup, handoff, and production workflows.
- Reduce manual production work through APIs, GitOps, automation, and agent-assisted workflows.
- Participate in on-call support, incident response, debugging, and follow-up remediation.
- Collaborate with platform, storage, networking, security, and workload teams to productionize infrastructure.
Requirements
- 8+ years of experience building or operating production infrastructure.
- Strong programming skills in Python, Go, or similar languages.
- Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation.
- Ability to troubleshoot distributed systems in production.
- Clear communication and ability to work across teams.
- BS/MS in Computer Science or equivalent experience.
Benefits
- Eligible for equity and benefits.
- Base salary is determined by location, experience, and comparable employee pay.
- Applications will be accepted at least until September 20, 2026.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
