1 day ago
Base Salary
$168k - $334k/yr
Responsibilities
- Build, implement, and support the operational and reliability aspects of large-scale Kubernetes clusters, including performance, monitoring, logging, and alerting.
- Define SLOs and SLIs, monitor error budgets, and streamline reliability reporting.
- Develop software tools, platforms, frameworks, and capacity-management processes to support services before launch.
- Maintain live services by measuring availability, latency, and overall system health.
- Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
- Automate and scale systems sustainably while driving improvements to reliability and engineering velocity.
- Lead triage and root-cause analysis for high-severity incidents, conduct blameless postmortems, and participate in on-call support.
Requirements
- BS in Computer Science or a related technical field, or equivalent experience.
- 8+ years of experience operating production services.
- Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
- Experience with infrastructure automation tools such as Terraform, Ansible, Chef, or Puppet.
- Proficiency in at least one high-level programming language such as Python or Go.
- In-depth knowledge of Linux, TCP/IP networking fundamentals, and cloud security standards.
- Knowledge of SRE principles including SLOs, SLIs, error budgets, and incident management.
- Experience building and operating observability stacks for monitoring, logging, and tracing.
- Preferred experience operating GPU-accelerated clusters with KubeVirt, applying generative AI to reduce operational toil, using workflow orchestration platforms, and operating AI inference workloads across the model-to-GPU stack.
Benefits
- Eligible for equity and benefits.
- Base salary is listed by location and level, with applications accepted at least until September 19, 2026.
Tech Stack
AnsibleApache AirflowAWSAzureChefGoGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPuppetPythonPyTorchSplunkTerraform
Categories
Site Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
