GrepJob
DigitalOcean

Staff Engineer, Inference Optimizations

DigitalOcean
Apply
about 2 hours ago
San Francisco, CA, USASenior / Staff+
H1B Sponsor

Base Salary

$191k - $239k/yr

Responsibilities

  • Lead the technical strategy for benchmarking and performance optimizations.
  • Engineer solutions for complex performance issues in AI inference.
  • Implement cutting-edge optimization techniques for Gen AI.
  • Act as a subject matter expert on modern GPU families and software stacks.
  • Develop and deploy state-of-the-art quantization techniques.
  • Provide technical mentorship through code and design reviews.
  • Collaborate with product management to translate hardware limits into product features.
  • Maintain a strong presence in GPU infrastructure and model performance optimization communities.

Requirements

  • 5+ years of experience in high-performance computing or AI infrastructure.
  • Deep familiarity with the Gen AI landscape and major model families.
  • Hands-on experience with attention-layer optimizations and parallelization strategies.
  • Comprehensive understanding of NVIDIA and AMD GPU architectures.
  • Extensive experience with open-source software projects.
  • Excellent system design skills related to low-level GPU programming.
  • Experience acting as a technical lead in cross-functional teams.
  • Deep understanding of GPU architectures and low-level programming.

Benefits

  • Career development resources including reimbursement for conferences and training.
  • Access to LinkedIn Learning's 10,000+ courses for continued growth.
  • Competitive benefits including Employee Assistance Program and flexible time off.
  • Potential for bonuses and equity compensation based on performance.

Categories

AI & MLData Engineering