1 day ago
Responsibilities
- Lead Day 2 operational readiness initiatives with NVIDIA Cloud Partners.
- Build continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters.
- Establish monitoring, alerting, dashboards, telemetry, and operational signals across infrastructure and AI workloads.
- Develop automated workflows for detecting, isolating, repairing, validating, and restoring unhealthy infrastructure.
- Manage GPU fleet lifecycle activities including driver and firmware updates, Kubernetes node maintenance, OS patching, upgrades, and configuration drift detection.
- Translate NVIDIA reference architectures into production operating practices, runbooks, automation, validation criteria, and operational standards.
- Define health signals, SLOs, metrics, acceptance criteria, and readiness mechanisms for infrastructure reliability.
- Create reusable operational frameworks, tooling, implementation guides, and playbooks for partner environments.
Requirements
- Bachelor’s, master’s, or Ph.D. degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience.
- At least 8 years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar production infrastructure roles.
- Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
- Deep knowledge of Kubernetes, containers, cluster scheduling, and large multi-node environment operations.
- Strong understanding of production observability, including metrics, logging, alerting, dashboards, health checks, and service-level agreement-driven operations.
- Experience automating infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
- Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers.
- Programming and automation experience with Python, Go, shell scripting, or similar languages.
- Preferred experience managing GPU or accelerated-computing infrastructure for AI training and inference workloads.
- Preferred experience with NVIDIA technologies, cloud partners or hyperscale providers, large-scale compute SLOs, infrastructure observability tools, and distributed AI workload failure modes.
Tech Stack
Categories
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
