Base Salary
$152k - $288k/yr
Responsibilities
- Own SRE solutions from design and implementation through operations and continuous improvement.
- Integrate reliability solutions with HPC schedulers, storage, and network fabrics.
- Standardize and automate provisioning using Infrastructure as Code and configuration management.
- Deliver services across globally distributed on-premises, AWS, GCP, and OCI environments.
- Design for failure using redundancy, failure domains, progressive delivery, and strict change control.
- Maintain uptime and quality of service through operational excellence, capacity planning, monitoring, and performance improvement.
- Participate in on-call rotations, incident reviews, root-cause analysis, and RCA report creation.
- Mentor engineers and influence technical direction through design reviews and architecture documentation.
Requirements
- Bachelor’s degree in Computer Science or a related technical field, or equivalent experience, plus 5+ years of professional experience building and supporting critical services.
- Experience supporting large-scale HPC clusters using Slurm or LSF, or Kubernetes clusters, including setup, tuning, and troubleshooting.
- Experience with CI/CD techniques and Infrastructure as Code for managing services.
- Strong experience building infrastructure platforms for automated host lifecycle management, fleet reliability and auto-healing, end-to-end observability, or data-driven operations using AIOps or ML-driven signals.
- Proficiency with monitoring, metrics, container management, and log collection tools.
- 5+ years of coding or scripting experience in at least two high-level languages such as Python, Go, Perl, or Ruby.
- Strong debugging, communication, documentation, mentoring, and technical collaboration skills.
- Published technical reliability, observability, or HPC work and production-scale open-source maintenance are preferred.
Benefits
- Competitive base salary of $152,000–$241,500 for Level 3 or $184,000–$287,500 for Level 4, plus equity and benefits.
- Applications will be accepted at least until September 8, 2026.
- NVIDIA is an equal opportunity employer committed to an inclusive work environment.
Tech Stack
Categories
H1B sponsorship
Sponsorship for this role is unconfirmed.
Nvidia’s past filings provide context. They do not guarantee sponsorship for a current opening.
H1B petition approvals by fiscal year
- FY 2026 · through Q2832
22 initial/new employment approvals · 355 changes of employer
- FY 20251,767
560 initial/new employment approvals · 415 changes of employer
- FY 20241,519
374 initial/new employment approvals · 415 changes of employer
Counts are petitions, including continuing employment and amendments, rather than unique hires. Source: USCIS
Job fields receiving H1B certifications
Partial fiscal year, through Q3
- Software Developers983 · 41%
- Electronics Engineers, Except Computer768 · 32%
- Architectural and Engineering Managers129 · 5%
- Computer and Information Research Scientists127 · 5%
- Sales Engineers77 · 3%
Job fields come from government occupation codes, rather than internal departments. Counts are certified Labor Condition Applications, not visa approvals or unique hires. Withdrawn and denied cases are excluded.
Source: US Department of LaborMatched government employer: NVIDIA CORPORATION
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
