DigitalOcean

Staff Engineer, Inference Optimizations

DigitalOcean
Apply
2 months ago

Base Salary

$191k - $239k/yr

Responsibilities

  • Lead technical strategy for benchmarking and performance optimization at inference-engine and GPU-kernel layers.
  • Optimize attention layers, memory and precision management, and parallel execution across multi-node GPU clusters.
  • Tune AMD AITER, composable kernel, assembly, FP8, BF16, FlashAttention, RMS Norm, and mixture-of-experts router kernels for large models.
  • Advise on NVIDIA and AMD hardware procurement and software integration.
  • Develop and deploy quantization techniques using FP8, INT8, and experimental FP4 to improve throughput while preserving accuracy.
  • Provide technical mentorship through code and design reviews without direct administrative management.
  • Partner with Product Management and TPMs to turn hardware capabilities and limits into shippable product features.
  • Contribute to and integrate open-source work within GPU infrastructure and model-performance optimization communities.

Requirements

  • 5+ years of experience in high-performance computing or AI infrastructure, including solving compute-utilization and memory-bandwidth bottlenecks.
  • Deep familiarity with generative AI and the architectures of major LLM, VLM, and LMM model families.
  • Hands-on experience with attention-layer optimization and parallelization across distributed GPU environments.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems, including CUDA and ROCm.
  • Extensive experience integrating, building with, and contributing to open-source software projects.
  • Strong system design skills involving low-level GPU programming, optimization, memory access patterns, and parallel execution.
  • Experience acting as a technical lead and driving design and delivery through cross-functional alignment and expert-level delegation.
  • Deep understanding of GPU architectures, including SMs, warp scheduling, and Tensor Cores.
  • Expert-level Triton or CUDA experience; Triton compiler contributions or custom CUDA kernels for major LLMs are valued.

Benefits

  • Remote role.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning with more than 10,000 courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Eligible employees may receive bonuses, equity grants upon hire, and access to an Employee Stock Purchase Program.

Categories

DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.

Contact me