
Staff Engineer, Inference Optimizations
DigitalOcean2 months ago
Base Salary
$191k - $239k/yr
Responsibilities
- Lead benchmarking and performance optimization strategy at the inference-engine and GPU-kernel layers.
- Optimize attention layers, memory and precision management, and parallel execution across multi-node GPU clusters.
- Tune AMD AITER, composable kernels, assembly kernels, and FP8/BF16 workloads for AMD MI355X GPUs.
- Identify kernel-fusion opportunities for Transformer components such as FlashAttention and RMS Norm.
- Tune expert-gateway router kernels for mixture-of-experts models including Qwen3-235B, DeepSeek V3, and GLM-5.
- Develop and deploy quantization techniques including FP8, INT8, and experimental FP4.
- Advise on modern GPU hardware procurement and software-stack integration.
- Provide technical mentorship through code and design reviews without direct administrative management.
- Partner with Product Management and TPMs to turn hardware capabilities and limits into shippable product features.
- Participate in GPU infrastructure and model-performance communities and contribute to open-source AI.
Requirements
- 5+ years of experience in high-performance computing or AI infrastructure.
- Proven experience solving compute-utilization and memory-bandwidth bottlenecks.
- Deep familiarity with LLM, VLM, and LMM architectures and major model families.
- Hands-on experience with attention-layer optimization and distributed GPU parallelization.
- Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems.
- Extensive experience integrating, building with, and contributing to open-source software.
- Strong system design skills involving low-level GPU programming, memory access patterns, and parallel execution.
- Experience acting as a technical lead and driving design and delivery through cross-functional alignment.
- Deep understanding of GPU architectures, SMs, warp scheduling, and Tensor Cores.
- Expert-level Triton or CUDA experience; Triton compiler contributions or custom CUDA kernels for major LLMs are particularly relevant.
Benefits
- Remote role.
- Conference, training, and education reimbursement.
- Access to LinkedIn Learning with 10,000+ courses.
- Employee Assistance Program and local employee meetups.
- Flexible time off policy.
- Eligible employees may receive equity grants upon hire and participate in the Employee Stock Purchase Program.
Tech Stack
DigitalOcean
Categories
About DigitalOcean
DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.