2 days ago
Base Salary
$108k - $196k/yr
Responsibilities
- Bring up, validate, tune, benchmark, and debug large-scale AI clusters and distributed training and inference workloads.
- Design and implement benchmarking tooling, automation, benchmark suites, acceptance criteria, and platform qualification workflows.
- Perform root-cause analysis and build tooling to detect, triage, and attribute node, fabric, and workload failures.
- Tune runtime settings, communication parameters, and deployment configurations across multi-GPU and multi-node systems.
- Provide data-driven recommendations based on profiling, benchmark results, and cluster characterization.
- Collaborate with framework, systems, and platform teams to improve resilience and performance.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related technical field, or equivalent experience.
- Experience developing software for AI, HPC, or systems-level applications.
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
- Experience debugging and scaling distributed systems and triaging AI applications from the application level toward the hardware.
- Experience operating workloads in scheduled, containerized cluster environments.
- Strong Python and C/C++ programming skills, along with analytical, debugging, communication, and collaboration abilities.
- Preferred experience includes NCCL, RDMA software such as IB verbs, UCX, and libfabric; InfiniBand or RoCE congestion debugging; MLPerf and benchmark or acceptance-test tooling; and resilience or failure-attribution systems.
Benefits
- Equity and benefits are provided.
- Base salary varies by location, experience, and comparable employee pay; the stated ranges are $108,000–$178,250 for Level 1 and $124,000–$195,500 for Level 2.
- Applications will be accepted at least until October 3, 2026.
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
