2 days ago
Austin, TX, USA +2 moreMid Level / Senior
H1B sponsor
Base Salary
$116k - $224k/yr
Responsibilities
- Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.
- Tune and benchmark AI pre-training, post-training, and inference workloads using NVIDIA AI software stacks.
- Perform root-cause analysis of failures in large distributed environments.
- Build resilience and failure-attribution tooling for node, fabric, and workload failures.
- Build and maintain benchmark suites, automation, acceptance criteria, regression gates, and platform qualification workflows.
- Tune runtime settings, communication parameters, and deployment configurations with framework, systems, and platform teams.
- Provide data-driven recommendations based on profiling, benchmark results, and cluster characterization.
Requirements
- Bachelor’s or master’s degree in computer science or a related technical field, or equivalent experience.
- 3+ years of experience developing software for AI, HPC, or systems-level applications.
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.
- Experience debugging and scaling distributed systems and triaging AI applications across the application-to-hardware stack.
- Experience operating workloads in scheduled, containerized cluster environments.
- Strong Python and C/C++ programming skills.
- Experience with NCCL, RDMA software, InfiniBand or RoCE congestion debugging, benchmark harnesses, cluster qualification tooling, MLPerf, performance jitter, or datacenter-scale resilience systems is preferred.
Benefits
- Base salary and equity eligibility, along with benefits.
- Salary ranges are provided for Level 2 and Level 3 positions.
- Applications will be accepted at least until October 3, 2026.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
