
Staff Engineer, Inference Optimizations
DigitalOceanabout 2 hours ago
Boston, MA, USASenior / Staff+
H1B Sponsor
Base Salary
$191k - $239k/yr
Responsibilities
- Lead the technical strategy for benchmarking and performance optimizations.
- Engineer solutions for complex performance issues in AI inference.
- Implement cutting-edge optimization techniques for Gen AI.
- Act as a subject matter expert on modern GPU families and software stacks.
- Develop and deploy state-of-the-art quantization techniques.
- Provide technical mentorship through code and design reviews.
- Collaborate with product management to translate hardware limits into product features.
- Maintain a strong presence in GPU infrastructure and model performance optimization communities.
Requirements
- 5+ years of experience in high-performance computing or AI infrastructure.
- Deep familiarity with the Gen AI landscape and major model families.
- Hands-on experience with attention-layer optimizations and parallelization strategies.
- Comprehensive understanding of NVIDIA and AMD GPU architectures.
- Extensive experience with open-source software projects.
- Excellent system design skills related to low-level GPU programming.
- Experience acting as a technical lead in cross-functional teams.
- Deep understanding of GPU architectures and low-level programming.
Benefits
- Career development resources including reimbursement for conferences and access to LinkedIn Learning.
- Competitive benefits including an Employee Assistance Program and flexible time off policy.
- Potential for bonuses and equity compensation based on performance.
Categories
AI & MLData Engineering