about 2 hours ago
Bellevue, WA, USA +2 moreMid Level / Senior
H1B Sponsor
Base Salary
$188k - $275k/yr
Responsibilities
- Build and maintain benchmarking workflows for measuring latency, throughput, and quality regressions.
- Benchmark the inference stack against customer workloads to identify performance gaps.
- Profile model-serving behavior to find bottlenecks in various systems.
- Drive optimization efforts for specific customer and product workloads.
- Design and run experiments on model-serving techniques with a focus on quality.
- Collaborate with platform engineers to productionize improvements.
- Produce clear technical writeups and recommendations for model configurations.
Requirements
- 4+ years of experience in machine learning, systems, or performance engineering.
- Strong programming skills in Python and experience in production environments.
- Experience running empirical evaluations and translating results into engineering decisions.
- Familiarity with LLM inference systems and model-serving stacks.
- Understanding of tradeoffs in latency, throughput, and quality regression analysis.
- Ability to work across model, systems, and product boundaries.
- Strong written communication skills.
Benefits
- 100% paid medical, dental, and vision insurance.
- Company-paid life insurance and disability insurance.
- Flexible Spending Account and Health Savings Account.
- Tuition reimbursement and employee stock purchase program.
- Mental wellness benefits and family-forming support.
- Paid parental leave and flexible childcare support.
- 401(k) with generous employer match and flexible PTO.
- Catered lunch and a casual work environment.
