DigitalOcean

Principal Engineer, Model Optimizations

DigitalOcean
Apply
14 hours ago

Base Salary

$250k - $312k/yr

Responsibilities

  • Set the technical strategy for model optimization across DigitalOcean’s accelerator fleet.
  • Own quantization strategy, calibration methodology, accuracy budgets, and evaluation gates.
  • Optimize modern model architectures through MoE routing, fused kernels, attention variants, and memory-movement improvements.
  • Lead speculative decoding development and acceptance-rate tuning.
  • Write and tune CUDA, Triton, CUTLASS, HIP, Composable Kernel, hipBLASLt, and AITER kernels when needed.
  • Define repeatable tensor, pipeline, expert, and attention parallelism layouts for models, GPUs, and traffic shapes.
  • Build benchmarking, regression, latency, throughput, and accuracy-validation infrastructure.
  • Drive AMD enablement and upstream contributions to vLLM, SGLang, and TensorRT-LLM.
  • Partner with NVIDIA and AMD engineering teams on pre-silicon enablement, early-access hardware, and roadmap feedback.
  • Set technical direction, mentor senior and staff engineers, and represent DigitalOcean with upstream communities, conferences, and customers.

Requirements

  • 12+ years of experience in performance-critical systems with substantial recent production experience optimizing LLM inference.
  • Deep understanding of GPU architecture and inference performance, including memory bandwidth, compute bounds, arithmetic intensity, scheduling overhead, and prefill versus decode.
  • Hands-on kernel-level experience with at least one vendor stack and the ability to work across CUDA/CUTLASS/Triton and ROCm/HIP/Composable Kernel.
  • Practical quantization expertise and judgment regarding production accuracy requirements.
  • Familiarity with the internals of at least one major serving engine: vLLM, SGLang, or TensorRT-LLM.
  • Strong Python and C++/CUDA skills and experience profiling with Nsight, rocprof, or equivalent tools.
  • Experience leading cross-functional efforts involving infrastructure, product, and customers.
  • Preferred qualifications include upstream contributions to vLLM, SGLang, TensorRT-LLM, PyTorch, Triton, or ROCm.
  • Preferred experience includes production enablement of new accelerator families, disaggregated prefill/decode serving, day-zero model enablement, and publications or patents in efficient inference, quantization, or GPU kernel design.

Benefits

  • Hybrid work arrangement.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses and career-development resources.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Potential bonus eligibility and equity compensation, including equity grants upon hire and an Employee Stock Purchase Program.
  • Equal-opportunity employment and benefits that may vary by location and local regulations.

Categories

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.

Contact me