5 days ago
Base Salary
$184k - $357k/yr
Responsibilities
- Architect scalable failure-attribution frameworks and high-fidelity flight recorders for EDA jobs across CPU, GPU, and fabric nodes.
- Build automated diagnostics correlating GPU XID errors, PCIe bus failures, CUDA memory exceptions, OOM kills, NUMA-related hangs, and system events.
- Implement low-overhead distributed logging and tracing for multi-node Slurm or Kubernetes clusters.
- Develop machine-learning-based heuristics and models to classify failures as hardware faults, software bugs, or environment issues.
- Define impending-failure signals and enable proactive job migration or checkpointing in collaboration with hardware and infrastructure teams.
Requirements
- BS, MS, or PhD in Computer Science or Electrical Engineering, or equivalent experience.
- 6+ years of experience in systems programming.
- Experience building automated root-cause-analysis pipelines for HPC or cloud-scale environments.
- Expertise in x86/ARM node-level metrics, including IPC, cache contention, NUMA imbalance, and hardware interrupts.
- Strong C++ and Python programming skills for building high-performance system-health daemons.
- Familiarity with Slurm, LSF, or Kubernetes and their job-lifecycle and signal-propagation mechanisms.
- Preferred: expert knowledge of the Linux kernel and interfaces such as /dev/mcelog, dmesg, and journald.
- Preferred: deep experience with NVIDIA DCGM and NVML for GPU health monitoring and state dumps.
- Preferred: experience with non-intrusive application-health and syscall-level monitoring.
- Preferred: experience with checkpoint/restore technologies such as CRIU in long-running EDA flows.
Benefits
- Hybrid work arrangement.
- Base salary range of 184,000 USD–287,500 USD for Level 4 or 224,000 USD–356,500 USD for Level 5.
- Eligible for equity and benefits.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
