
Principal Engineer, Model Optimizations
DigitalOcean14 hours ago
Base Salary
$250k - $312k/yr
Responsibilities
- Set the technical strategy for model optimization across DigitalOcean’s accelerator fleet.
- Own quantization strategy, calibration methodology, accuracy budgets, and evaluation gates.
- Optimize modern model architectures through MoE routing, fused kernels, attention variants, and memory-movement improvements.
- Lead speculative decoding development and acceptance-rate tuning.
- Write and tune CUDA, Triton, CUTLASS, HIP, Composable Kernel, hipBLASLt, and AITER kernels when needed.
- Define repeatable tensor, pipeline, expert, and attention parallelism layouts for models, GPUs, and traffic shapes.
- Build benchmarking, regression, latency, throughput, and accuracy-validation infrastructure.
- Drive AMD enablement and upstream contributions to vLLM, SGLang, and TensorRT-LLM.
- Partner with NVIDIA and AMD engineering teams on pre-silicon enablement, early-access hardware, and roadmap feedback.
- Set technical direction, mentor senior and staff engineers, and represent DigitalOcean with upstream communities, conferences, and customers.
Requirements
- 12+ years of experience in performance-critical systems with substantial recent production experience optimizing LLM inference.
- Deep understanding of GPU architecture and inference performance, including memory bandwidth, compute bounds, arithmetic intensity, scheduling overhead, and prefill versus decode.
- Hands-on kernel-level experience with at least one vendor stack and the ability to work across CUDA/CUTLASS/Triton and ROCm/HIP/Composable Kernel.
- Practical quantization expertise and judgment regarding production accuracy requirements.
- Familiarity with the internals of at least one major serving engine: vLLM, SGLang, or TensorRT-LLM.
- Strong Python and C++/CUDA skills and experience profiling with Nsight, rocprof, or equivalent tools.
- Experience leading cross-functional efforts involving infrastructure, product, and customers.
- Preferred qualifications include upstream contributions to vLLM, SGLang, TensorRT-LLM, PyTorch, Triton, or ROCm.
- Preferred experience includes production enablement of new accelerator families, disaggregated prefill/decode serving, day-zero model enablement, and publications or patents in efficient inference, quantization, or GPU kernel design.
Benefits
- Hybrid work arrangement.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning courses and career-development resources.
- Employee Assistance Program, local employee meetups, and flexible time off.
- Potential bonus eligibility and equity compensation, including equity grants upon hire and an Employee Stock Purchase Program.
- Equal-opportunity employment and benefits that may vary by location and local regulations.
Categories
About DigitalOcean
DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.