1 year ago
Base Salary
$150k - $230k/yr
Responsibilities
- Build inference services with smart batching and caching.
- Optimize kernels, tokenization, and model graphs.
- Evaluate tradeoffs among vLLM, TensorRT LLM, and Triton.
- Implement autoscaling and admission control with clear SLOs.
- Own performance dashboards and capacity planning.
- Specialize in low-latency, high-throughput inference for OCR and multimodal models across single-tenant and multi-tenant environments.
Requirements
- At least 3 years of experience in performance engineering or ML systems.
- Strong Python skills.
- C++ or CUDA exposure.
- Experience with GPU profiling and model serving.
- Experience reducing p95 latency and cost in production ML systems is preferred.
Benefits
- Five days per week in the San Francisco office.
- Competitive base salary plus equity and a performance-based bonus.
- Relocation assistance for Bay Area moves.
- Daily meal stipend.
- Medical, vision, and dental coverage.
