11 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms.
- Lead bare-metal provisioning and lifecycle management, including PXE boot, DHCP, DNS, operating-system provisioning, hardware validation, and fleet automation.
- Develop automation, self-service capabilities, observability solutions, and infrastructure API integrations using Python or Go.
- Define and operate SLOs, SLIs, error budgets, alerting, incident-response practices, and reliability improvements.
- Lead complex incident investigations, corrective actions, and blameless postmortems.
- Partner with infrastructure, security, hardware, data-center, and application teams on global platform initiatives.
- Participate in an on-call rotation.
Requirements
- Bachelor of Science in Computer Science, Engineering, a related technical field, or equivalent experience.
- 10+ years of experience operating production infrastructure or platform services.
- Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, Linux systems, and distributed-system troubleshooting.
- Experience deploying and operating bare-metal infrastructure in data-center environments, including provisioning, networking, operating-system lifecycle management, and hardware automation.
- Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet.
- Understanding of TCP/IP networking and infrastructure security.
- Strong SRE and observability experience with SLIs, SLOs, error budgets, incident management, monitoring, logging, and tracing.
- Preferred experience with HPC, AI, GPU-accelerated, GPU-enabled Kubernetes or KubeVirt infrastructure, VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.
- Preferred experience applying generative AI or agentic workflows to infrastructure diagnostics and incident resolution.
- Experience building secure operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, and workflow orchestration.
- Demonstrated delivery of complex, high-impact infrastructure projects.
Benefits
- Participation in an on-call rotation is required.
- The role supports global engineering workloads and collaborates across infrastructure, security, hardware, data-center, and application teams.
Tech Stack
Categories
DevOpsSite Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
