about 3 hours ago
San Francisco, CA, USA or Sunnyvale, CA, USAMid Level / Senior
H1B Sponsor
Base Salary
$250k - $300k/yr
Responsibilities
- Optimize current inference techniques for production use.
- Design and enhance serving architectures for AI models.
- Profile and analyze performance issues down to the kernel level.
- Adapt optimization methods for various ML models, focusing on large language models.
- Tune deployments for latency, throughput, and cost efficiency.
- Collaborate with customer engineering teams to transition workloads to production.
- Develop software features around the inference stack using general-purpose languages.
- Conduct rapid experiments to refine goals into actionable specs.
- Manage end-to-end delivery from experimentation to production optimization.
- Navigate ambiguity and make informed decisions on tradeoffs.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field.
- Hands-on experience with production code in Python or C++.
- Familiarity with optimizing LLMs for high throughput and low latency.
- Experience with modern LLM serving frameworks like vLLM or SGLang.
- Understanding of GPU architecture and behavior.
- Interest and experience with large language models.
- Knowledge of AI/ML pipelines and model deployment.
- Strong communication skills for technical discussions.
Benefits
- Competitive compensation and equity packages.
- Restricted Stock Units.
- Paid time off and holidays.
- Comprehensive health, dental, and vision insurance.
- Employer contributions to HSA accounts.
- Paid parental leave and life insurance.
- Professional development and tuition reimbursement.
- Mental health and wellness support.
- Commuter benefits and cell phone stipend.
- 401(k) retirement plan with company match.
- Volunteer time off and global travel insurance.
- Daily meals allowance and additional location-specific perks.
Tech Stack
Categories
AI & MLData Engineering
