about 3 hours ago
Palo Alto, CA, USASenior
Base Salary
$195k - $262k/yr
Responsibilities
- Own optimization work for specific model families and customer endpoints.
- Run engine comparisons and recommend serving configurations for workloads.
- Debug model quality or performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for various performance metrics.
- Deploy and extend inference engines like vLLM and TensorRT-LLM.
- Build and productionize model-compression workflows.
- Implement advanced decoding and serving techniques.
- Build reproducible benchmark harnesses for performance evaluation.
- Collaborate with engineers to diagnose performance bottlenecks.
- Write design docs, performance reports, and technical explanations.
Requirements
- Strong Python and PyTorch engineering skills.
- Experience deploying or optimizing LLM or high-throughput transformer systems.
- Knowledge of modern inference stacks like Triton Inference Server.
- Understanding of transformer inference bottlenecks.
- Ability to analyze latency, throughput, and cost tradeoffs.
- Strong communication skills for collaboration across teams.
Benefits
- 100% company-paid medical, dental, and vision coverage.
- 401(k) plan with up to 4% company match.
- 20 weeks paid parental leave for primary caregivers.
- Remote work reimbursement up to $85/month.
- Company-paid short-term, long-term, and life insurance.