Responsibilities
- Lead major incidents end to end, including triage, cross-team coordination, decision-making, and executive communication across global time zones.
- Set technical direction for SRE initiatives that improve reliability, scalability, and developer efficiency across enterprise systems.
- Design, build, and operate distributed systems and Kubernetes-based cloud-native infrastructure.
- Build automation for incident detection, triage, communication, and remediation, including self-healing systems.
- Improve observability and signal quality to detect issues earlier, reduce alert noise, and eliminate dependence on user-reported issues.
- Lead root-cause analysis and convert incident learnings into systemic fixes, automation, and prevention mechanisms.
- Apply LLMs, anomaly detection, and signal correlation to incident triage, summarization, and decision support.
- Champion AI-assisted engineering practices, including coding agents and LLM-powered tooling.
- Partner with Cloud, Platform, Security, and AI/ML teams to define SLOs and error budgets and influence architecture.
- Mentor engineers, conduct design and code reviews, and strengthen the reliability culture in the India organization.
Requirements
- 10+ years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Incident Management roles.
- BS or MS degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- Experience acting as an Incident Commander or leading major incident response in complex, high-availability environments.
- Deep understanding of distributed systems, monitoring, reliability engineering, SLIs/SLOs, error budgets, capacity planning, and graceful degradation.
- Strong proficiency in at least one programming language such as Python, Go, or Java.
- Hands-on expertise with AWS, Azure, or GCP and container technologies including Docker and Kubernetes.
- Experience with infrastructure-as-code tooling such as Terraform, AWS CDK, or CloudFormation and CI/CD pipelines.
- Strong Linux/Unix and networking fundamentals and expertise with OpenTelemetry, Prometheus, and Grafana.
- Working knowledge of relational databases such as PostgreSQL and MySQL, including SQL, indexing, and query optimization.
- Familiarity with AI/ML concepts such as LLMs, anomaly detection, and data-driven operations applied to operational workflows.
- Excellent written and verbal communication skills, including executive briefings during high-pressure incidents and influence with senior technical stakeholders.
- Preferred experience building AI-enabled incident-management or automation platforms, applying intelligent triage, RCA generation, and signal correlation, and reducing MTTD and MTTR.
- Preferred experience scaling reliability across distributed teams, follow-the-sun operations, and AI/ML training and inference infrastructure across multiple regions.
- Contributions to open-source infrastructure or observability projects, conference talks, or the broader SRE community are valued.
Benefits
- NVIDIA offers highly competitive salaries and a comprehensive benefits package for employees and their families.
- The role is based in India and is an individual-contributor position with mentoring and technical leadership responsibilities.
Tech Stack
Categories
H1B sponsorship
Sponsorship for this role is unconfirmed.
Nvidia’s past filings provide context. They do not guarantee sponsorship for a current opening.
H1B petition approvals by fiscal year
- FY 2026 · through Q2832
22 initial/new employment approvals · 355 changes of employer
- FY 20251,767
560 initial/new employment approvals · 415 changes of employer
- FY 20241,519
374 initial/new employment approvals · 415 changes of employer
Counts are petitions, including continuing employment and amendments, rather than unique hires. Source: USCIS
Job fields receiving H1B certifications
Partial fiscal year, through Q3
- Software Developers983 · 41%
- Electronics Engineers, Except Computer768 · 32%
- Architectural and Engineering Managers129 · 5%
- Computer and Information Research Scientists127 · 5%
- Sales Engineers77 · 3%
Job fields come from government occupation codes, rather than internal departments. Counts are certified Labor Condition Applications, not visa approvals or unique hires. Withdrawn and denied cases are excluded.
Source: US Department of LaborMatched government employer: NVIDIA CORPORATION
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
