1 day ago
Base Salary
$152k - $242k/yr
Responsibilities
- Design, implement, and support operational and reliability aspects of large-scale observability and telemetry collection platforms.
- Improve services across their full lifecycle, from inception and design through deployment, operation, and refinement.
- Develop software tools, platforms, and frameworks; provide system design consulting, capacity management, and launch reviews.
- Measure and monitor service availability, latency, and overall system health after launch.
- Scale systems through automation and drive improvements to reliability and engineering velocity.
- Participate in sustainable incident response, blameless postmortems, and an on-call rotation for production systems.
Requirements
- Bachelor’s degree in Computer Science or a related coding-focused technical field, such as physics or mathematics, or equivalent experience.
- At least 5 years of experience with infrastructure automation, distributed systems design, and tools for operating large-scale private or public cloud systems in production.
- At least 5 years of experience delivering foundational infrastructure and observability platforms.
- Experience with one or more of Python, Go, Perl, or Ruby.
- In-depth knowledge of Linux, networking, and containers.
- Interest in analyzing and fixing large-scale distributed systems, with systematic problem-solving, communication, ownership, and debugging skills.
- Experience running large private or public cloud systems based on Kubernetes, OpenStack, and Docker, and experience running Grafana, OpenTelemetry, Prometheus, or similar observability tools.
Benefits
- Eligible for equity and benefits.
- Base salary range is $152,000-$241,500, determined by location, experience, and comparable employee pay.
- Applications will be accepted at least until September 15, 2026.
- The role includes participation in an on-call rotation.
Tech Stack
Categories
DevOpsSite Reliability
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
