Nvidia

Senior Staff Site Reliability Engineer

Nvidia
Apply
5 hours ago
Bengaluru, IndiaStaff+

Responsibilities

  • Define the architecture and technical roadmap for a scalable enterprise AI runtime platform.
  • Design Kubernetes-based systems for deploying, operating, and scaling AI applications, inference services, and databases across cloud and on-premises environments.
  • Build control-plane services, APIs, operators, and automation for workload provisioning, configuration, upgrades, and recovery.
  • Develop GPU scheduling, autoscaling, load balancing, and rate-limiting capabilities for runtime services.
  • Improve the performance, availability, and developer experience of large-scale AI inference services.
  • Build and automate relational and vector database services, including provisioning, scaling, backup, and failover.
  • Establish secure application lifecycle-management patterns across cloud and on-premises platforms.
  • Develop observability tools for monitoring, profiling, and debugging applications, GPU resources, inference workloads, and databases.
  • Lead technical initiatives across multiple functions, mentor engineers, and establish platform standards.

Requirements

  • Bachelor’s, master’s, or doctoral degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 8+ years of software engineering experience building distributed systems, cloud infrastructure, database platforms, or large-scale backend services.
  • Strong programming skills in Python, Go, C++, or Java with experience delivering production-grade systems.
  • Experience designing scalable, highly available Kubernetes-based platforms and leading technical strategy across teams.
  • Experience building control planes, platform APIs, Kubernetes operators, or workload lifecycle-management systems.
  • Experience developing high-performance services for AI inference or other low-latency workloads.
  • Practical knowledge of relational or vector databases, including availability, replication, query optimization, and performance tuning.
  • Experience with GitOps, CI/CD, observability, and cloud-native security practices.
  • Preferred experience with self-service platforms, inference-serving frameworks, GPU-aware scheduling, model-performance optimization, vector databases, GPU-accelerated query engines, distributed data platforms, full AI application lifecycle support, or relevant open-source contributions.

Benefits

  • NVIDIA offers competitive salaries and a comprehensive benefits package for employees and their families.
  • The company supports a diverse and equal-opportunity work environment.

Categories

BackendDevOpsSite Reliability
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me