Nvidia

Senior Software Engineer, SRE and Production Engineering - DGX Cloud

Nvidia
Apply
2 days ago
Santa Clara, CA, USASenior
H1B sponsor

Base Salary

$152k - $288k/yr

Responsibilities

  • Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management.
  • Develop tools using BMC and Redfish interfaces to monitor hardware health, manage server state, and support recovery workflows.
  • Operate and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs across cloud partner and on-premises environments.
  • Diagnose failures across servers, DPUs, GPUs, CPUs, networking, Linux, and Kubernetes, and automate recurring detection and repair.
  • Define validation and handoff criteria for safely and consistently placing new capacity into production.
  • Participate in on-call duties, incident response, root-cause analysis, and permanent remediation.
  • Collaborate with hardware, networking, platform, data center operations, and partner teams on cross-functional issues.

Requirements

  • 5+ years of experience building software for or operating production infrastructure, including substantial hands-on bare-metal experience.
  • Strong Go or Python skills and a record of delivering production automation and services.
  • Direct experience with BMC and Redfish for server provisioning, health inspection, power management, or fault diagnosis.
  • Practical experience with NVIDIA GPU hardware such as NVL72 systems and BlueField-3 or newer DPUs.
  • Experience with Linux, firmware and driver management, network boot, and the server lifecycle from provisioning through repair.
  • Experience managing production reliability through on-call duties, incident response, observability, and durable solutions.
  • Ability to debug failures across hardware, host operating systems, networking, and distributed services.
  • Clear communication and demonstrated ownership of problems spanning multiple teams.
  • BS or MS in Computer Science or equivalent experience in a practical setting.
  • Preferred experience includes BlueField DPU mode, host-to-DPU connectivity, NVLink, InfiniBand, Spectrum-X, GPU cluster performance validation, rack-scale bringup, firmware upgrades, hardware replacement, Kubernetes, GitOps, Argo CD, SLOs, and fleet-wide automation.

Benefits

  • Base salary and equity eligibility are provided, along with benefits.
  • Applications will be accepted at least until October 3, 2026.

Tech Stack

Categories

DevOpsSite Reliability
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me