Nvidia

Senior Software Engineer, DGX Cloud AI Infrastructure

Nvidia
Apply
2 days ago
Austin, TX, USA +2 moreSenior
H1B sponsor

Base Salary

$184k - $357k/yr

Responsibilities

  • Lead bring-up, validation, debugging, and qualification of large-scale AI clusters, infrastructure, and end-to-end workloads.
  • Tune and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo/Megatron, TensorRT-LLM, and NVIDIA AI software stacks.
  • Profile and optimize workload performance across compute, memory, networking, and communication layers using Nsight Systems, NCCL tests, and custom microbenchmarks.
  • Analyze distributed LLM scaling using data, tensor, pipeline, and expert parallelism across GPU clusters.
  • Own root-cause analysis for failures, hangs, performance regressions, and topology sensitivity in distributed environments.
  • Build resilience and failure-attribution systems for node, fabric, and workload failures at cluster scale.
  • Create benchmark suites, automation, acceptance criteria, regression gates, and platform qualification workflows.
  • Tune runtime, communication, and deployment configurations with framework, systems, and platform teams.
  • Deliver data-driven recommendations from profiling, benchmarking, and cluster characterization.
  • Mentor engineers and drive technical standards across the performance and infrastructure organization.

Requirements

  • Bachelor’s or master’s degree in computer science or a related technical field, or equivalent experience.
  • 8+ years of experience developing software infrastructure for large-scale AI or HPC systems, including technical leadership.
  • Expertise debugging and triaging AI applications from the application layer through the hardware layer.
  • Deep hands-on experience with NCCL, CUDA-aware distributed execution, and multi-GPU and multi-node workload debugging.
  • Track record architecting, debugging, and scaling large-scale distributed systems.
  • Expert-level Python and C/C++ programming skills.
  • Experience operating workloads in scheduled, containerized cluster environments.
  • Strong analytical, debugging, communication, and cross-team influence skills.
  • Preferred experience with large-scale AI workload optimization, RDMA software, GPU fabrics and topology, benchmark harnesses, qualification tooling, and resilience or failure-attribution systems.

Benefits

  • Base salary is listed by level and location, experience, and comparable employee pay; employees are also eligible for equity and benefits.
  • Applications will be accepted at least until October 3, 2026.
  • This posting is for an existing vacancy.

Tech Stack

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me