
Staff Software Engineer: AI Inference Data Plane
DigitalOcean3 days ago
Base Salary
$167k - $209k/yr
Responsibilities
- Lead the end-to-end design, development, and delivery of high-scale data-plane components for generative AI inference.
- Architect resilient, highly available, multi-tenant inference cloud services.
- Optimize distributed inference using parallelism, KV-cache optimization, smart routing, and caching techniques.
- Build Kubernetes-native distributed inference systems supporting prefill/decode disaggregation, cache-aware routing, autoscaling, and MoE expert parallelism.
- Solve inference-specific distributed-systems challenges involving load balancing, flow control, tenant fairness, latency, and KV-cache transfer.
- Contribute to llm-d, vLLM, and related inference gateway open-source communities.
- Mentor junior engineers and provide technical leadership across engineering, product, and customer-facing teams.
- Operate critical services, define SLOs, and use observability tools to maintain platform health.
Requirements
- Hands-on experience hosting large language or multimodal models with inference engines such as vLLM, SGLang, or TensorRT.
- Familiarity with distributed inference frameworks including llm-d, NVIDIA Dynamo, or Ray Serve.
- Experience with inference-engine internals such as continuous batching, paged attention, and prefix caching.
- Understanding of cluster-scale inference, KV-cache locality, prefill/decode disaggregation, and fast cross-pod KV transfer.
- Expert-level proficiency in GoLang or Python and familiarity with gRPC.
- Knowledge of LLM architectures and optimization techniques such as quantization.
- Experience shipping customer-facing software and operating critical, high-scale services.
- Experience integrating and building with open-source software.
- Merged contributions to vLLM, llm-d, SGLang, or similar projects is strongly preferred.
Benefits
- Remote work arrangement.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning’s 10,000+ courses.
- Employee Assistance Program, local employee meetups, and flexible time off.
- Potential bonus eligibility and equity compensation, including equity grants upon hire and an Employee Stock Purchase Program.
Tech Stack
Categories
About DigitalOcean
DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.