2 days ago
Base Salary
$184k - $357k/yr
Responsibilities
- Develop infrastructure software and tools for large-scale AI pre-training, post-training, and inference.
- Develop and optimize tools and libraries that improve infrastructure efficiency and resiliency.
- Co-design and implement APIs integrating with NVIDIA resiliency stacks.
- Enhance infrastructure and products supporting NVIDIA AI platforms.
- Define reliability metrics to track and improve system and service reliability.
- Debug, perform root-cause analysis, and triage failures from the application level through the hardware level.
Requirements
- At least 8 years of experience developing software infrastructure for large-scale AI systems.
- Bachelor’s degree or higher in Computer Science or a related technical field, or equivalent experience.
- Strong debugging, analysis, and AI application failure-triage skills.
- Experience with observability platforms for monitoring and logging, such as ELK, Prometheus, and Loki.
- Proven experience building and scaling large-scale distributed systems.
- Experience with AI training and inference infrastructure services.
- Proficiency in Python, C, C++, and scripting languages.
- Experience with software engineering practices including test development, defensive programming, version control, and CI.
- Preferred experience with large-scale clusters, observability and telemetry stacks, RDMA software, datacenter-scale failure analysis, NCCL, IB verbs, UCX, libfabric, PyTorch, TensorFlow, JAX, and Ray.
- Strong communication, collaboration, problem-solving, and analytical skills.
Benefits
- Equity and benefits are provided.
- The base salary range is $184,000-$287,500 for Level 4 and $224,000-$356,500 for Level 5.
- The application will be accepted at least until October 3, 2026.
- The position is an existing vacancy.
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
