Nvidia

Senior Staff Site Reliability Engineer - Compute Core Engineering

Nvidia
Apply
6 days ago
Remote, Worldwide or Santa Clara, CA, USAStaff+
H1B Sponsor

Base Salary

$200k - $322k/yr

Responsibilities

  • Lead initiatives to transform the Compute Core team’s architecture and build new on-premises and cloud service offerings.
  • Design, scale, and deploy globally reliable core infrastructure services including DNS, NTP/PTP, DHCP, and LDAP.
  • Implement automation, monitoring, high availability, capacity planning, and lifecycle management for infrastructure services.
  • Define and track efficiency metrics and drive software and hardware optimizations using technologies such as SR-IOV and DPU.
  • Use eBPF and XDP for observability and DDoS mitigation.
  • Analyze system and capacity data, develop enterprise-wide capacity plans, and coordinate infrastructure changes with management.
  • Develop and maintain tools for data collection, analysis, visualization, reporting, alerting, and monitoring.
  • Collaborate with leadership, engineers, program managers, and product managers to develop IT products and services.

Requirements

  • Bachelor’s degree in Engineering, Computer Science, Mathematics, or a related field, or equivalent experience.
  • 12+ years of experience in compute platform engineering with a focus on automation.
  • Experience designing and deploying containerization architectures and distributed-systems infrastructure.
  • Experience evaluating application architectures and identifying containerization opportunities to improve scalability, reliability, and efficiency.
  • Strong analytical skills and experience defining and tracking key performance metrics.
  • Experience developing data-analysis and performance-profiling tools, with Terraform and configuration-management tools.
  • Proficiency in Go and/or Python.
  • Proficiency with Linux and kernel internals.
  • Experience operating large bare-metal build-infrastructure environments.
  • Understanding of network protocols and architectures including VLAN, VxLAN, SDN, BGP, and Anycast.
  • Preferred experience with containers, DNS and LDAP services at scale, microservices architecture, infrastructure as code, configuration management, and security tools.

Benefits

  • Base salary range of 200,000 USD to 322,000 USD.
  • Eligible for equity and benefits.
  • Applications accepted at least until September 1, 2026.
  • NVIDIA offers an inclusive work environment and equal opportunity employment.

Categories

DevOpsSite Reliability
Nvidia

About Nvidia

10,000+ employees

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Contact me