DigitalOcean

Principal Engineer, Inference Memory and Storage Systems

DigitalOcean
Apply
13 hours ago

Base Salary

$250k - $312k/yr

Responsibilities

  • Define the technical vision and roadmap for a unified multi-tier memory and storage layer supporting large-scale LLM inference.
  • Architect integrations with vLLM, SGLang, and TensorRT-LLM for KV-cache offload, reuse, eviction, and cross-node sharing.
  • Design protocols for disaggregated prefill/decode, peer-to-peer KV-cache transfer, and cache-aware routing.
  • Own cache admission, eviction, pluggable backend, and metrics strategies across memory and storage tiers.
  • Partner with GPU infrastructure, networking, and platform teams to use GPUDirect, RDMA, NVMe-oF, and NVLink for low-latency access.
  • Model and validate the impact of cache hit ratio, prefix caching, and tiering on cost and performance.
  • Set technical direction, lead design reviews, mentor senior and staff engineers, and sponsor follow-on initiatives.
  • Contribute to open source, deliver conference talks, and conduct customer-facing technical deep dives.

Requirements

  • 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure, including production service delivery and operations.
  • Deep understanding of GPU HBM, host DRAM, NVMe, remote storage, and tiered memory architectures.
  • Experience with distributed caching or key-value systems optimized for low latency and high concurrency.
  • Hands-on experience with networked I/O, RDMA, NVMe-oF, NVLink-class technologies, and aggregated or disaggregated AI-cluster topologies.
  • Strong systems programming skills in C, C++, Go, Rust, or Python, with the ability to read and modify serving-engine internals.
  • Experience profiling and optimizing CPU, GPU, memory, and network performance using metrics such as TTFT, ITL, and throughput.
  • Strong written and verbal communication skills and experience leading cross-functional efforts.
  • Preferred experience contributing to vLLM, SGLang, llm-d, NVIDIA Dynamo, LMCache, or similar inference infrastructure projects.
  • Preferred experience designing unified KV or object models across GPU, host, SSD, and cloud storage tiers.
  • Preferred familiarity with Kubernetes-based GPU orchestration, DRA, MIG/MPS partitioning, and gateway or inference-extension routing.
  • Publications or patents in LLM systems, memory-disaggregated architectures, RDMA data planes, or ML caching systems are a bonus.

Benefits

  • Hybrid work arrangement.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning with more than 10,000 courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Potential bonus and equity compensation, including equity grants upon hire and participation in the Employee Stock Purchase Program.
  • Competitive benefits that may vary according to local regulations and preferences.
DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.

Contact me