Nvidia

Distinguished Engineer, Production Engineering, Cluster Management

Nvidia
Apply
1 day ago
Santa Clara, CA, USAStaff+
H1B Sponsor

Base Salary

$320k - $489k/yr

Responsibilities

  • Define long-range technical strategy for operating DGX Cloud clusters across local data centers, hyperscalers, and NeoCloud environments.
  • Set architectural direction and operating standards for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability.
  • Guide cross-organizational investments that improve production readiness, operational safety, performance, and coordination.
  • Build workflows, interfaces, automation, APIs, and readiness gates connecting Kubernetes services, providers, hardware, on-premises infrastructure, and bare-metal environments.
  • Implement operating approaches that reduce manual work, clarify ownership, improve traceability, and increase release safety.
  • Partner with platform, hardware, provider, service-owner, and Production Engineering teams to convert recurring friction into durable software, process, and interface improvements.
  • Raise engineering standards for operability, resilience, scalability, and performance through implementation leadership, architecture reviews, and technical standards.
  • Lead technical decisions and cross-team efforts from concept through production and drive adoption across organizational boundaries.

Requirements

  • Bachelor’s, master’s, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience.
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments.
  • Company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms.
  • Track record defining operating models, architectural direction, and engineering standards across multiple technical domains and organizations.
  • Track record leading large cross-team technical efforts from concept through production and delivering measurable outcomes.
  • Deep experience with Kubernetes-based production systems, infrastructure automation, or distributed systems operations.
  • Strong software engineering skills in Python, Go, or similar low-level programming languages.
  • Deep understanding of distributed systems, Linux, networking, containers, and production reliability concerns.
  • Experience creating operational workflows, APIs, service interfaces, or automation frameworks that become standard production practices.
  • Strong architectural judgment and experience simplifying complex operational problems through reusable software and durable technical strategy.
  • Preferred experience establishing foundations, standards, architectures, APIs, or workflows for large heterogeneous infrastructure environments spanning multiple platforms or providers.
  • Preferred experience improving production readiness, restoration, runtime safety, or release quality for large-scale infrastructure.
  • Ability to set strategy, review architecture at scale, write code, and drive adoption across organizational boundaries.

Benefits

  • Base salary range of 320,000 USD to 488,750 USD, determined by location, experience, and comparable employee pay.
  • Eligible for equity and benefits.
  • Hybrid work arrangement.
  • Applications accepted at least until September 7, 2026.
  • This posting is for an existing vacancy.

Categories

Nvidia

About Nvidia

10,000+ employees

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Contact me