Nvidia

Senior Deep Learning Engineer – Autonomous Vehicles

Nvidia
Apply
1 day ago
Boulder, CO, USA or Santa Clara, CA, USASenior
H1B sponsor

Base Salary

$224k - $357k/yr

Responsibilities

  • Build, scale, and harden deep learning infrastructure libraries and frameworks for training on multi-thousand-GPU clusters.
  • Improve training-stack efficiency across data loaders, distributed training, scheduling, and performance monitoring.
  • Develop robust training pipelines and libraries for massive video datasets and rapid experimentation.
  • Collaborate with researchers, model engineers, and internal platform teams to reduce stalls and improve training availability.
  • Own orchestration libraries, distributed training frameworks, and fault-resilient training systems.
  • Partner with leadership to scale infrastructure with GPU capacity and dataset growth while maintaining developer efficiency and stability.

Requirements

  • Bachelor’s, master’s, or PhD in Computer Science, Electrical/Computer Engineering, or a related field, or equivalent experience.
  • 12+ years of professional experience building and scaling high-performance distributed systems, ideally in ML, HPC, or large-scale data infrastructure.
  • Extensive knowledge of deep learning frameworks, especially PyTorch, large-scale training with DDP/FSDP, NCCL, tensor parallelism, pipeline parallelism, and performance profiling.
  • Strong systems background including datacenter networking, parallel filesystems, storage systems, and schedulers.
  • Proficiency in Python and C++ with experience developing production-grade libraries, orchestration layers, and automation tools.
  • Ability to collaborate with ML researchers, infrastructure engineers, and product leads and translate requirements into robust systems.
  • Preferred experience scaling GPU training clusters with more than 1,000 GPUs.
  • Preferred contributions to open-source ML systems libraries such as PyTorch, NCCL, FSDP, schedulers, or storage clients.
  • Preferred expertise in fault resilience, high availability, elastic training, and large-scale observability.
  • Preferred familiarity with reinforcement learning at scale, particularly for simulation-heavy workloads.
  • Hands-on technical leadership experience establishing guidelines for ML systems engineering.

Benefits

  • Eligible for equity and benefits.
  • Base salary is determined by location, experience, and pay for similar positions, with a stated range of 224,000 USD to 356,500 USD.
  • Applications accepted at least until September 12, 2026.
Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me