4 days ago
Base Salary
$248k - $391k/yr
Responsibilities
- Define architecture, service tiers, SLAs, and automated cluster lifecycles for a global compute platform running OpenShift and KubeVirt.
- Build automated remediation pipelines, hardware watchdogs, and telemetry for pre-release rack-scale GPU systems and internal AI inference infrastructure.
- Drive capacity planning and scale strategies including public cloud bursting, hardware dogfooding, and evaluation of alternative compute architectures such as ARM.
- Design self-service platform architectures, APIs, and Terraform/OpenTofu providers to promote standardized internal platforms.
- Lead migrations of large legacy workloads, including long-running VDI environments, into modern Kubernetes orchestration.
- Influence technical direction across autonomous engineering teams and establish operational maturity through self-service, auto-remediation, and strict SLAs.
Requirements
- Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
- 15+ years of experience in compute platform engineering, site reliability, or systems architecture with extensive automation at massive scale.
- Deep expertise in Kubernetes architecture and virtualization architectures that operate VMs inside Kubernetes, specifically KubeVirt and OpenShift.
- In-depth knowledge of GPUs and high-speed backplane networking, including mitigation of hardware failures, silent data corruption, and large-scale anomalies.
- Experience operating global bare-metal, virtualized, and cloud environments with a unified GitOps posture using ArgoCD or similar tools.
- Proficiency in Go and/or Python and expert-level infrastructure-as-code development using Terraform and configuration management.
- Strong leadership and influence skills for driving technical direction across highly autonomous teams.
- Preferred experience managing pre-release hardware in production, advanced storage migrations and protocols, multi-cloud deployment, and building Day 2 operational maturity from existing foundations.
Benefits
- Hybrid work arrangement.
- Eligibility for equity and benefits.
- Applications accepted at least until September 3, 2026.
- The posting is for an existing vacancy.
Tech Stack
Categories
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
