DigitalOcean

Senior Engineer II, Inference Engine - Serving Engine

DigitalOcean
Apply
3 days ago

Base Salary

$167k - $209k/yr

Responsibilities

  • Lead the end-to-end design, development, and delivery of data plane components hosting large generative AI models.
  • Architect high-scale, multi-tenant AI inference services with strong availability and resiliency.
  • Implement and optimize distributed inference using tensor/data parallelism, KV-cache optimization, and smart routing.
  • Collaborate with product managers, customer-facing teams, and engineering teams on technical roadmaps.
  • Coach and mentor junior engineers.
  • Operate critical services, use observability tools, and define SLOs for platform health.

Requirements

  • Strong experience with distributed systems, microservices, messaging systems, databases, and infrastructure as code.
  • Hands-on experience hosting large language or multimodal models with inference engines such as vLLM, SGLang, or Modular.
  • Familiarity with distributed inference serving frameworks including llm-d, NVIDIA Dynamo, or Ray Serve.
  • Understanding of GPU-level optimization and interconnect technologies such as NVLink, XGMI, or RoCE.
  • Knowledge of LLM architectures and optimization techniques including continuous batching and quantization.
  • Expert-level proficiency in Go or Python and familiarity with gRPC.
  • Experience shipping customer-facing software and operating critical services at high scale.
  • Experience integrating and building with open-source software.

Benefits

  • Remote work arrangement.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Potential bonus and equity compensation, including eligible equity grants and an Employee Stock Purchase Program.

Tech Stack

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.

Contact me