3 days ago
Base Salary
$152k - $288k/yr
Responsibilities
- Design and build automation platforms for provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
- Develop monitoring, health-management, remediation, and recovery systems for EDA compute environments.
- Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
- Build services and workflows integrating workload schedulers, infrastructure management systems, and observability platforms.
- Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and restore unhealthy systems.
- Collaborate with EDA, infrastructure, networking, storage, and hardware engineering teams on scalable chip-design workload solutions.
- Participate in incident response, root-cause analysis, capacity planning, and continuous production-service improvement.
Requirements
- 5+ years of software engineering or infrastructure engineering experience supporting large-scale production systems.
- BS in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
- Strong programming experience in Go or Python with knowledge of data structures, algorithms, testing, and software design.
- Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
- Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
- Experience with infrastructure automation, software deployment, observability, and operational recovery.
- Strong communication, cross-team collaboration, systematic problem-solving, and ownership skills.
- Preferred experience with large-scale EDA or high-performance computing infrastructure, Linux, GPU and CPU server architecture, networking, storage, and bare-metal lifecycle management.
- Preferred experience with Slurm, LSF, Kubernetes, Bright Cluster Manager, EDA applications, license-management systems, high-throughput batch workloads, or semiconductor design workflows.
- Preferred experience building health checks, break-fix remediation, firmware and operating-system upgrade workflows, node-provisioning systems, and multi-data-center or heterogeneous-hardware infrastructure.
Benefits
- Eligible for equity and benefits.
- Base salary varies by location and experience, with posted ranges of $152,000-$241,500 for Level 3 and $184,000-$287,500 for Level 4.
- Applications accepted at least until September 4, 2026.
About Nvidia
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
