DigitalOcean

Staff Engineer, Inference Optimizations

DigitalOcean
Apply
2 months ago
Boston, MA, USAStaff+
H1B sponsor

Base Salary

$191k - $239k/yr

Responsibilities

  • Lead technical strategy for benchmarking and performance optimization at inference-engine and GPU-kernel layers.
  • Optimize attention layers, memory and precision management, and parallelization across multi-node GPU clusters.
  • Tune AITER, composable kernel, assembly, FP8, and BF16 workloads for AMD MI355X GPUs.
  • Identify kernel-fusion opportunities for Transformer layers such as FlashAttention and RMS Norm.
  • Tune expert-gateway router kernels for mixture-of-experts models including Qwen3-235B, DeepSeek V3, and GLM-5.
  • Advise on modern NVIDIA and AMD GPU hardware procurement and software integration.
  • Develop and deploy quantization techniques using FP8, INT8, and experimental FP4.
  • Provide technical leadership through code and design reviews without direct people-management responsibilities.
  • Partner with Product Management and Technical Program Managers to turn hardware capabilities into shippable product features.
  • Participate in GPU infrastructure and model-performance communities and contribute to open-source AI.

Requirements

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Proven experience solving compute-utilization and memory-bandwidth bottlenecks.
  • Deep familiarity with generative AI, including LLM, VLM, LMM, and major model-family architectures.
  • Hands-on experience with attention-layer optimization and parallelization across distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems.
  • Extensive experience integrating, building with, and contributing to open-source software.
  • Excellent system design skills involving low-level GPU programming, optimization, memory access patterns, and parallel execution.
  • Experience acting as a technical lead and driving design and delivery through cross-functional alignment.
  • Deep understanding of GPU architectures, including SMs, warp scheduling, and Tensor Cores.
  • Expert-level Triton or CUDA experience; Triton compiler contributions or custom CUDA kernels for major LLMs are valued.

Benefits

  • Remote role.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Eligible employees may receive equity compensation and participate in the Employee Stock Purchase Program.
  • Bonus eligibility based on company and individual performance.

Categories

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.

Contact me