Nvidia

Senior Engineer, NCX

Nvidia
Apply
2 days ago
Remote, WorldwideSenior
H1B Sponsor

Responsibilities

  • Lead NVIDIA Cloud Partner Day 2 operational readiness efforts after initial cluster deployment and activation.
  • Develop continuous validation for GPU, CPU, storage, and network health across large-scale AI clusters.
  • Establish telemetry, monitoring, alerting, dashboards, health checks, and operational signals across compute, networking, storage, Kubernetes, and AI workloads.
  • Build automated workflows to detect, isolate, drain, repair, validate, and return unhealthy infrastructure to service.
  • Manage GPU fleet lifecycle activities including driver and firmware administration, Kubernetes node maintenance, OS patching, upgrades, configuration management, and drift detection.
  • Translate NVIDIA reference architectures and partner requirements into production practices, validation criteria, runbooks, automation, and operational standards.
  • Define health signals, SLOs, metrics, acceptance criteria, and validation mechanisms for infrastructure reliability and service readiness.
  • Create reusable tooling, implementation guides, runbooks, playbooks, and reference implementations for multiple partner environments.

Requirements

  • Bachelor’s, master’s, or Ph.D. in Computer Science, Computer or Electrical Engineering, or a related technical field, or equivalent experience.
  • At least 8 years of experience in infrastructure engineering, Site Reliability Engineering, DevOps, cloud platform engineering, systems engineering, or similar roles supporting large-scale production environments.
  • Strong experience operating Linux-based distributed systems and cloud infrastructure in production.
  • Deep understanding of Kubernetes, containers, cluster scheduling, and large multi-node environment operations.
  • Strong production observability experience spanning metrics, logging, alerting, dashboards, health checks, and service-level agreements.
  • Experience automating infrastructure lifecycle management, failure detection, remediation, upgrades, and configuration management.
  • Strong networking fundamentals and experience troubleshooting distributed systems across compute, network, and storage layers.
  • Programming and automation experience with Python, Go, shell scripting, or similar languages.
  • Preferred experience managing GPU or accelerated computing infrastructure for AI training and inference workloads.
  • Preferred experience with NVIDIA DGX/HGX systems, CUDA, NVLink/NVSwitch, NVIDIA networking, InfiniBand, RoCE, GPU Operator, or Network Operator.
  • Preferred experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or large service-provider infrastructures while operating SLOs.
  • Preferred knowledge of Prometheus, Grafana, OpenTelemetry, Alertmanager, scalable telemetry pipelines, distributed AI workload failure modes, and production operating models.

Tech Stack

GoGrafanaKubernetesLinuxPrometheusPython

Categories

DevOpsSite Reliability
Nvidia

About Nvidia

10,000+ employees

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Contact me