Nvidia

Senior Storage Software Engineer - DGX Cloud

Nvidia
Apply
2 days ago
Santa Clara, CA, USASenior
H1B sponsor

Base Salary

$224k - $431k/yr

Responsibilities

  • Contribute production code, fixes, and features to open-source parallel and distributed file systems and distributed object storage.
  • Serve as a hands-on technical lead by writing and reviewing code, examining kernel and storage-system source, and making technical delivery decisions.
  • Triage and root-cause storage issues across GPU clusters, including I/O and metadata performance, data corruption, and recovery problems.
  • Run scale tests, benchmarks, recovery drills, and architecture validation against measurable performance and durability targets.
  • Define storage configuration, tuning, and operational best practices and help operators and internal customers apply them.
  • Collaborate with training, inference, accelerated-computing, SRE, operations, networking, security, cloud, neocloud, and vendor teams.

Requirements

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • More than 12 years of direct storage software engineering experience, including extensive work with a high-performance parallel or distributed file system at multi-petabyte scale.
  • Contributions to open-source distributed or parallel file-system projects and demonstrated hands-on production engineering experience.
  • Experience diagnosing and resolving storage problems in large GPU or HPC clusters, including I/O and metadata performance analysis.
  • Strong proficiency in at least one of C, C++, Rust, or Go, plus proficiency in Python.
  • Familiarity with Linux kernel storage and networking stacks, including block layer, RDMA, RoCE, InfiniBand, NVMe, page cache, VFS, and multipath.
  • Understanding of object storage such as S3 or Swift and block storage such as NVMe-oF and iSCSI.
  • Strong written and verbal communication skills and ability to explain technical trade-offs to engineers, SREs, vendors, and customers.
  • Comfort operating in a 24/7 production environment where storage incidents affect GPU availability, with a security-first approach.
  • Preferred qualifications include maintaining widely used public projects; AI training or inference storage at very large GPU scale; kernel or file-system development; metadata scalability; data placement; failure recovery; HSM or equivalent experience; Kubernetes and CSI driver development; and SPDK, libfabric, or FUSE performance optimization.

Benefits

  • Eligible for equity and benefits.
  • Base salary range is $224,000-$356,500 for Level 5 or $272,000-$431,250 for Level 6, depending on location, experience, and comparable employee pay.
  • Applications will be accepted at least until September 24, 2026.
  • This posting is for an existing vacancy.

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me