2 months ago
Remote, United StatesStaff+
Base Salary
$185k - $280k/yr
Responsibilities
- Optimize and deploy high-performance LLM inference pipelines across data center, edge, and embedded platforms.
- Own, extend, and tune vLLM, TensorRT-LLM, llama.cpp, and QAIRT inference runtimes.
- Develop custom CUDA kernels and adapt inference runtimes for constrained and embedded deployment environments.
- Implement and evaluate INT8, INT4, FP4, FP8, mixed-precision, AWQ, and GPTQ quantization strategies.
- Optimize KV cache performance through paging, prefix caching, and cache-aware memory layout design.
- Design and tune batching, continuous batching, and speculative decoding to improve latency and throughput.
- Reduce inference costs, memory pressure, and production latency while improving tokens-per-second performance.
Requirements
- Proven experience optimizing machine learning inference performance in production.
- Deep understanding of GPU architecture and memory hierarchies.
- Hands-on experience with CUDA and low-level performance tuning.
- Experience deploying machine learning models beyond research environments.
- Experience with vLLM, TensorRT-LLM, llama.cpp, and QAIRT inference engines.
- Experience with CUDA kernel development and profiling, quantization techniques, KV cache optimization, memory layout design, batching, speculative decoding, and continuous batching.
Benefits
- Annual bonus opportunity.
- Medical, dental, vision, life, and disability insurance coverage.
- Paid time off and paid holidays.
- Company contribution to the RRSP.
- Equity awards for certain positions and levels.
- Remote and/or hybrid work available depending on the position.