
Principal Engineer, Rack-Scale GPU Architecture
DigitalOcean13 hours ago
Base Salary
$250k - $312k/yr
Responsibilities
- Own the end-to-end reference architecture for rack-scale inference systems, from fabric design through customer scheduling.
- Define NVLink-domain partitioning, allocation, isolation, multi-node provisioning, and fabric-manager configuration.
- Architect and tune scale-out RDMA fabrics using RoCEv2 or InfiniBand, rail-optimized topologies, congestion control, GPUDirect RDMA, and NCCL.
- Make Kubernetes topology-aware through DRA, LeaderWorkerSet, gang scheduling, Topology Manager, NUMA alignment, and GPU and Network Operators.
- Drive multi-node disaggregated inference, including prefill and decode placement and KV-cache movement.
- Own rack-scale failure semantics, diagnostics, drain and repair workflows, degraded-domain scheduling, and firmware and driver lifecycle.
- Establish rack validation, burn-in, bandwidth and latency characterization, NCCL benchmarking, inference benchmarking, and production acceptance gates.
- Partner with data center engineering on power, liquid cooling, rack density, deployability, and cost constraints.
- Work with NVIDIA and ODM partners on pre-silicon enablement, hardware bringup, and roadmap feedback.
- Set technical direction, mentor senior and staff engineers, and represent DigitalOcean with upstream communities, conferences, and customers.
Requirements
- 12+ years of large-scale compute, HPC, or GPU infrastructure engineering experience with hands-on production responsibility.
- Deep understanding of GPU interconnect and system topology, including NVLink, NVSwitch, NVLink domains, PCIe, and NUMA.
- Strong RDMA networking experience with RoCEv2 or InfiniBand, GPUDirect RDMA, congestion control, lossless fabric tuning, and rail-optimized cluster topologies.
- Practical NCCL expertise involving algorithm and protocol selection, topology detection, and collective-performance debugging.
- Substantial Kubernetes experience for accelerated workloads, including device plugins, DRA, scheduling extensions, and operators.
- Strong systems programming and automation skills in Go and Python, with comfort at the driver, firmware, and BMC layers.
- Experience operating fleets where hardware failure is routine, including diagnostics, automation, and blast-radius management.
- Excellent written and verbal communication and experience leading architecture across hardware, networking, platform, and product teams.
- Preferred experience bringing up GB200 or GB300 NVL72 or comparable rack-scale platforms in production.
- Preferred experience with multi-node inference, liquid-cooled high-density racks, heterogeneous accelerator fleets, and relevant open-source projects such as Kubernetes SIG-Node, NVIDIA GPU or Network Operator, DRA drivers, Kueue, LeaderWorkerSet, or NCCL.
Benefits
- Hybrid work arrangement.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning courses.
- Employee Assistance Program, local employee meetups, and flexible time off.
- Potential performance-based bonus and equity compensation, including eligible equity grants and Employee Stock Purchase Program participation.
Tech Stack
Categories
About DigitalOcean
DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.