Nvidia

Senior Production Engineer - DGX Cloud

Nvidia
Apply
1 day ago
Remote, Switzerland or Zürich, SwitzerlandSenior

Responsibilities

  • Build and operate production software, automation, and tooling for control-plane services, model deployments, and inference and agentic workloads.
  • Improve reliability, availability, routing, capacity management, observability, rollout safety, and recovery for inference platforms and services.
  • Deploy, configure, validate, upgrade, and recover services consistently across environments using infrastructure as code and GitOps.
  • Automate service enablement, model releases, handoff, deprecation, and ongoing operational workflows.
  • Define and instrument SLIs and SLOs, use error budgets, and improve visibility into production health.
  • Participate in on-call and incident response, troubleshoot failures across routing, runtimes, software, and infrastructure, and create durable automated fixes.
  • Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams.

Requirements

  • 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation.
  • Strong programming skills in Python, Go, or a comparable language, with experience developing production-operations tools.
  • Experience with infrastructure as code, configuration management, or GitOps and repeatable service-deployment automation.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals.
  • Understanding of SLIs, SLOs, error budgets, incident response, and operational-toil reduction.
  • Experience using metrics, logs, and traces to understand system behavior and improve reliability.
  • Clear technical communication and ability to work across engineering teams.
  • BS/MS in Computer Science or equivalent practical experience.
  • Preferred experience with vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, GPU performance analysis, Kubernetes operators, controllers, workload orchestration, fleet management, self-healing automation, Terraform, Argo CD, policy validation, safe deployment and rollback systems, AI tools and agents, or production AI inference and agentic workloads.

Categories

DevOpsSite Reliability
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me