19 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Support production Kubernetes services as part of a 24/7 Production Engineering team, including split-weekend shifts.
- Automate operations and reduce manual tasks across large-scale production environments.
- Administer Kubernetes, systems, security monitoring, and large-scale clusters to maintain service availability, integrity, reliability, and SLAs.
- Use alerts, alarms, and observability tools to monitor services, detect issues, prevent incidents, and respond to failures.
- Analyze logs, metrics, and system behavior; troubleshoot issues, lead root cause analysis, and implement effective resolutions.
- Lead incident management calls and coordinate subject matter experts and service owners through detection, escalation, and resolution.
Requirements
- 7+ years of experience administering large-scale production Kubernetes systems in high-availability Internet, cloud, or data center environments, with strong preference for on-premises experience.
- Bachelor of Science degree in Computer Science, Engineering, or Mathematics, or equivalent experience.
- Advanced hands-on experience with Kubernetes, SLURM, and large-scale cluster management.
- Familiarity with GPU/DPU hardware and high-performance computing cluster environments.
- Strong Linux systems administration, DNS, DHCP, iptables, routing, firewall, and core Linux networking experience on large-scale bare-metal infrastructure.
- Experience with CI/CD tools such as Jenkins and Argo CD.
- Python, Golang, or Rust scripting/programming experience is preferred but not required.
- Strong communication skills and the ability to present persuasively to cross-functional groups.
Benefits
- Membership in a 24/7 Production Engineering team supporting production services.
- Split-weekend shift flexibility is required.
- Work includes large-scale cloud, data center, and potentially on-premises infrastructure environments.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
