2 days ago
Base Salary
$184k - $357k/yr
Responsibilities
- Build and operate production software, automation, and tooling for control plane services, model deployments, inference workloads, and agentic workloads.
- Improve reliability, availability, routing, capacity management, observability, rollout safety, recovery, and service health across DGX Cloud environments.
- Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently.
- Create workflows for service enablement, model releases, handoff, deprecation, and ongoing operations while automating repetitive manual work.
- Define and instrument SLIs and SLOs, use error budgets, and make production health visible to partner teams.
- Participate in on-call and incident response, diagnose failures, and convert recurring issues into automation and durable fixes.
- Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams.
Requirements
- 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation.
- Strong programming skills in Python, Go, or a comparable language, with experience developing production operations tools.
- Experience with infrastructure as code, configuration management, or GitOps and repeatable service deployment automation.
- Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals.
- Understanding of SRE principles including SLIs, SLOs, error budgets, incident response, and reducing operational toil.
- Experience using metrics, logs, and traces to understand system behavior and improve reliability.
- Clear technical communication and ability to work across engineering teams.
- BS/MS in Computer Science or equivalent experience.
- Preferred experience with vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, GPU performance analysis, Kubernetes operators, controllers, workload orchestration, fleet management, self-healing automation, Terraform, Argo CD, policy validation, safe deployment and rollback systems, AI tools and agents, or production AI inference and agentic workloads.
Benefits
- Base salary is provided in two bands: 184,000 USD–287,500 USD for Level 4 and 224,000 USD–356,500 USD for Level 5.
- Eligible for equity and benefits.
- The role involves multi-cloud and on-premises deployments across AWS, Azure, Google Cloud, partner cloud environments, and on-call operations.
- Applications are accepted at least until October 6, 2026.
Categories
DevOpsSite Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
