DigitalOcean

Principal Engineer, Rack-Scale GPU Architecture

DigitalOcean
Apply
13 hours ago

Base Salary

$250k - $312k/yr

Responsibilities

  • Own the end-to-end reference architecture for rack-scale inference systems, from fabric design through customer scheduling.
  • Define NVLink-domain partitioning, allocation, isolation, multi-node provisioning, and fabric-manager configuration.
  • Architect and tune scale-out RDMA fabrics using RoCEv2 or InfiniBand, rail-optimized topologies, congestion control, GPUDirect RDMA, and NCCL.
  • Make Kubernetes topology-aware through DRA, LeaderWorkerSet, gang scheduling, Topology Manager, NUMA alignment, and GPU and Network Operators.
  • Drive multi-node disaggregated inference, including prefill and decode placement and KV-cache movement.
  • Own rack-scale failure semantics, diagnostics, drain and repair workflows, degraded-domain scheduling, and firmware and driver lifecycle.
  • Establish rack validation, burn-in, bandwidth and latency characterization, NCCL benchmarking, inference benchmarking, and production acceptance gates.
  • Partner with data center engineering on power, liquid cooling, rack density, deployability, and cost constraints.
  • Work with NVIDIA and ODM partners on pre-silicon enablement, hardware bringup, and roadmap feedback.
  • Set technical direction, mentor senior and staff engineers, and represent DigitalOcean with upstream communities, conferences, and customers.

Requirements

  • 12+ years of large-scale compute, HPC, or GPU infrastructure engineering experience with hands-on production responsibility.
  • Deep understanding of GPU interconnect and system topology, including NVLink, NVSwitch, NVLink domains, PCIe, and NUMA.
  • Strong RDMA networking experience with RoCEv2 or InfiniBand, GPUDirect RDMA, congestion control, lossless fabric tuning, and rail-optimized cluster topologies.
  • Practical NCCL expertise involving algorithm and protocol selection, topology detection, and collective-performance debugging.
  • Substantial Kubernetes experience for accelerated workloads, including device plugins, DRA, scheduling extensions, and operators.
  • Strong systems programming and automation skills in Go and Python, with comfort at the driver, firmware, and BMC layers.
  • Experience operating fleets where hardware failure is routine, including diagnostics, automation, and blast-radius management.
  • Excellent written and verbal communication and experience leading architecture across hardware, networking, platform, and product teams.
  • Preferred experience bringing up GB200 or GB300 NVL72 or comparable rack-scale platforms in production.
  • Preferred experience with multi-node inference, liquid-cooled high-density racks, heterogeneous accelerator fleets, and relevant open-source projects such as Kubernetes SIG-Node, NVIDIA GPU or Network Operator, DRA drivers, Kueue, LeaderWorkerSet, or NCCL.

Benefits

  • Hybrid work arrangement.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Potential performance-based bonus and equity compensation, including eligible equity grants and Employee Stock Purchase Program participation.

Categories

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.

Contact me