
Staff Engineer, Inference Optimizations
DigitalOcean3 days ago
Base Salary
$191k - $239k/yr
Responsibilities
- Lead technical strategy for inference benchmarking and performance optimization at the inference-engine and GPU-kernel layers.
- Optimize attention layers, memory and precision management, and parallel execution across multi-node GPU clusters.
- Tune AITER, composable and assembly kernels, Transformer kernels, FlashAttention, RMS Norm, and expert-gateway router kernels for large MoE models.
- Develop and deploy FP8, INT8, and experimental FP4 quantization techniques to improve throughput while preserving accuracy.
- Serve as a subject matter expert on NVIDIA and AMD GPU architectures, hardware procurement, and associated software ecosystems.
- Provide technical mentorship through high-quality code and design reviews without direct people-management responsibilities.
- Partner with Product Management and TPMs to translate hardware constraints into shippable product features.
- Contribute to and integrate open-source AI and GPU infrastructure projects and participate in relevant technical communities.
Requirements
- At least 5 years of experience in high-performance computing or AI infrastructure, including solving compute-utilization and memory-bandwidth bottlenecks.
- Deep familiarity with LLM, VLM, and LMM architectures and the requirements of major model families.
- Hands-on experience with attention-layer optimization and distributed GPU parallelization.
- Comprehensive understanding of NVIDIA and AMD GPU architectures and their software ecosystems.
- Extensive experience integrating, building with, and contributing to open-source software.
- Strong system-design skills involving low-level GPU programming, memory-access patterns, and parallel execution.
- Experience serving as a technical lead and driving design and delivery through cross-functional influence.
- Deep understanding of GPU SMs, warp scheduling, and Tensor Cores.
- Expert-level experience with Triton or CUDA; Triton compiler contributions or custom CUDA kernels for major LLMs are valued.
Benefits
- Hybrid work arrangement.
- Competitive benefits, including an Employee Assistance Program, local employee meetups, and flexible time off, with specific offerings varying by location.
- Reimbursement for relevant conferences, training, and education.
- Access to LinkedIn Learning courses.
- Potential bonus based on company and individual performance.
- Equity compensation for eligible employees, including grants upon hire and participation in the Employee Stock Purchase Program.
Tech Stack
Assembly
Categories
About DigitalOcean
DigitalOcean is the AI-Native Cloud purpose-built for the inference and agentic era. Its five-layer integrated platform—spanning GPU and CPU infrastructure, core cloud, inference, data, and managed agent orchestration—is open throughout with no vendor lock-in, giving builders everything they need to start fast, scale production AI workloads, and improve unit economics. More than 650,000 customers and millions of developers globally trust DigitalOcean to build, ship, and scale their applications.