1 day ago
Base Salary
$248k - $397k/yr
Responsibilities
- Define and drive the long-term technical vision, architecture, and roadmap for reliability across NVIDIA’s AI Platform Runtime and related enterprise systems.
- Architect highly available, resilient, secure, and scalable distributed platforms for AI-driven products and services.
- Design AI agents, AI skills, and intelligent automation for platform operations, incident response, troubleshooting, and remediation.
- Establish reliability standards covering service-level objectives, error budgets, capacity models, resilience patterns, and operational readiness.
- Lead cross-functional programs addressing availability, scalability, performance, security, systemic risk, and developer productivity.
- Advance observability using OpenTelemetry, metrics, logs, traces, profiling, analytics, and automated anomaly detection.
- Partner with Cloud, Platform, Security, Networking, and AI/ML organizations on architectural and investment decisions.
- Provide technical leadership during critical incidents and turn incident lessons into durable improvements.
- Develop reference architectures, shared platform capabilities, and automation frameworks for adoption across engineering organizations.
- Influence and mentor senior engineers and technical leaders.
Requirements
- 15+ years of experience in Site Reliability Engineering, Platform Engineering, Distributed Systems, Cloud Architecture, or related infrastructure engineering roles.
- BS or MS degree in Computer Science or a related technical field involving significant software development, or equivalent experience.
- Experience setting technical strategy and leading large-scale engineering initiatives across multiple teams or organizations.
- Deep expertise in distributed systems architecture, networking, Linux, Kubernetes, and public cloud platforms such as AWS, Azure, or GCP.
- Strong proficiency in one or more of Python, Go, TypeScript, JavaScript, or Java, with experience building production-grade automation and platform software.
- Extensive experience with infrastructure-as-code and platform automation technologies such as Terraform, Crossplane, AWS CDK, or AWS CloudFormation.
- Deep understanding of observability at scale, including OpenTelemetry and modern metrics, logging, tracing, profiling, and analytics platforms.
- Expertise in service-level objectives, error budgets, capacity planning, fault tolerance, disaster recovery, incident management, and blameless postmortems.
- Track record of measurable improvements in availability, performance, operational efficiency, or engineering productivity in complex environments.
- Strong technical judgment, communication, collaboration, and ability to influence senior stakeholders.
- Preferred experience with large-scale AI/ML platforms, GPU infrastructure, inference systems, training environments, or high-performance computing platforms.
- Preferred hands-on experience building AI agents, agentic workflows, or intelligent infrastructure automation.
- Preferred experience applying machine learning, analytics, or advanced automation to capacity management, anomaly detection, incident response, or autonomous remediation.
- Preferred recognized technical leadership through patents, publications, open-source contributions, industry standards, conference presentations, or widely adopted internal platforms.
Benefits
- Hybrid work arrangement.
- Eligible for equity and benefits.
- Applications accepted at least until September 12, 2026.
- This posting is for an existing vacancy.
Categories
Site Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
