
Principal Engineer, Inference Memory and Storage Systems
DigitalOcean13 hours ago
Base Salary
$250k - $312k/yr
Responsibilities
- Define the technical vision and roadmap for a unified multi-tier memory and storage layer supporting large-scale LLM inference.
- Architect integrations with vLLM, SGLang, and TensorRT-LLM for KV-cache offload, reuse, eviction, and cross-node sharing.
- Design protocols for disaggregated prefill/decode, peer-to-peer KV-cache transfer, and cache-aware routing.
- Own cache admission, eviction, pluggable backend, and metrics strategies across memory and storage tiers.
- Partner with GPU infrastructure, networking, and platform teams to use GPUDirect, RDMA, NVMe-oF, and NVLink for low-latency access.
- Model and validate the impact of cache hit ratio, prefix caching, and tiering on cost and performance.
- Set technical direction, lead design reviews, mentor senior and staff engineers, and sponsor follow-on initiatives.
- Contribute to open source, deliver conference talks, and conduct customer-facing technical deep dives.
Requirements
- 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure, including production service delivery and operations.
- Deep understanding of GPU HBM, host DRAM, NVMe, remote storage, and tiered memory architectures.
- Experience with distributed caching or key-value systems optimized for low latency and high concurrency.
- Hands-on experience with networked I/O, RDMA, NVMe-oF, NVLink-class technologies, and aggregated or disaggregated AI-cluster topologies.
- Strong systems programming skills in C, C++, Go, Rust, or Python, with the ability to read and modify serving-engine internals.
- Experience profiling and optimizing CPU, GPU, memory, and network performance using metrics such as TTFT, ITL, and throughput.
- Strong written and verbal communication skills and experience leading cross-functional efforts.
- Preferred experience contributing to vLLM, SGLang, llm-d, NVIDIA Dynamo, LMCache, or similar inference infrastructure projects.
- Preferred experience designing unified KV or object models across GPU, host, SSD, and cloud storage tiers.
- Preferred familiarity with Kubernetes-based GPU orchestration, DRA, MIG/MPS partitioning, and gateway or inference-extension routing.
- Publications or patents in LLM systems, memory-disaggregated architectures, RDMA data planes, or ML caching systems are a bonus.
Benefits
- Hybrid work arrangement.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning with more than 10,000 courses.
- Employee Assistance Program, local employee meetups, and flexible time off.
- Potential bonus and equity compensation, including equity grants upon hire and participation in the Employee Stock Purchase Program.
- Competitive benefits that may vary according to local regulations and preferences.
Tech Stack
Categories
About DigitalOcean
DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.