Nvidia

Senior Software Engineer - Cluster Networking

Nvidia
Apply
1 day ago
Durham, NC, USA +3 moreSenior
H1B sponsor

Base Salary

$184k - $357k/yr

Responsibilities

  • Own and evolve Kubernetes networking architecture for GPU clusters operating at multi-thousand-node scale.
  • Design, operate, and scale CNI data planes, overlay meshes, VPN topologies, and gateways connecting control and data planes.
  • Design and operate L7 gateways, load balancers, and tunnels using Envoy and Cloudflare.
  • Identify and eliminate scale limitations involving packet loss, control-plane saturation, IP address exhaustion, and large-scale failure modes.
  • Build scale-test environments and validation suites to detect networking regressions before production workloads are affected.
  • Diagnose distributed networking problems across Kubernetes, Slurm, training jobs, hosts, routes, packet marking, and tunnels.
  • Partner with cloud and neocloud providers on network topology, requirements, and capabilities.
  • Provide senior technical judgment and networking expertise to a distributed engineering team.

Requirements

  • BS or MS in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
  • At least 6 years of professional experience in systems, network, or infrastructure software engineering.
  • Deep knowledge of Kubernetes networking architecture and CNI standards; production experience with Calico is strongly preferred.
  • Experience designing and maintaining mesh and VPN networking topologies using Tailscale, WireGuard, or equivalent technologies.
  • Strong Linux networking fundamentals, including routing, netfilter, iptables/nftables, packet marking, network namespaces, and container-runtime interactions.
  • Ability to debug distributed network problems at scale using packet capture, tracing, and multi-host root-cause analysis.
  • Proficiency in Go, Python, C, or a comparable systems programming language.
  • Clear written and verbal communication and ability to collaborate across multiple time zones.
  • Experience with massive-scale Kubernetes topologies, high-performance AI or HPC fabrics such as InfiniBand, RoCE, or RDMA, upstream contributions to relevant projects, multi-cloud and on-premises networking, or Slurm and other HPC schedulers is valued.

Benefits

  • Equity and benefits are provided.
  • Base salary is determined by location, experience, and comparable employee pay.
  • Applications are accepted at least until September 19, 2026.

Tech Stack

AmbassadorCCloudflareGoKubernetesLinuxPython

Categories

Nvidia

About Nvidia

10,000+ employees

Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.

Contact me