DigitalOcean

Staff Engineer, Inference Optimizations

DigitalOcean
Apply
2 months ago

Base Salary

$191k - $239k/yr

Responsibilities

  • Lead technical strategy for inference-engine and GPU-kernel benchmarking and performance optimization.
  • Optimize attention layers, memory and precision management, and parallelization across multi-node GPU clusters.
  • Tune AMD AITER, CK, ASK, FP8, and BF16 workloads for AMD MI355X GPUs.
  • Identify kernel-fusion opportunities for GLM-5 Transformer layers, including FlashAttention and RMS Norm.
  • Tune expert-gateway router kernels for mixture-of-experts models such as Qwen3-235B, DeepSeek V3, and GLM-5.
  • Advise on modern NVIDIA and AMD GPU hardware, software integration, and hardware procurement.
  • Develop and deploy quantization techniques using FP8, INT8, and FP4 to improve throughput while preserving accuracy.
  • Guide the technical roadmap for the high-performance inference fleet and act as a technical force multiplier.
  • Provide technical mentorship through high-quality code and design reviews without direct management responsibilities.
  • Collaborate with Product Management and TPMs to translate hardware capabilities and limits into shippable product features.
  • Participate in GPU infrastructure and model-performance communities and contribute to open-source AI.

Requirements

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Proven experience solving compute-utilization and memory-bandwidth bottlenecks.
  • Deep familiarity with LLM, VLM, and LMM landscapes and major model-family architectures.
  • Hands-on experience with attention-layer optimization and parallelization across distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems.
  • Extensive experience integrating, building with, and contributing to open-source software projects.
  • Strong system design skills involving low-level GPU programming, memory access patterns, and parallel execution.
  • Experience serving as a technical lead and driving design and delivery through cross-functional alignment.
  • Deep understanding of GPU architectures, including SMs, warp scheduling, and Tensor Cores.
  • Expert-level Triton or CUDA experience; Triton compiler contributions or custom CUDA kernels for major LLMs are relevant examples.

Benefits

  • Remote role.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Potential bonus and equity compensation, including equity grants upon hire and an Employee Stock Purchase Program.

Categories

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.

Contact me