
LLM Inference Engineer
Hippocratic AI11 months ago
Responsibilities
- Design and implement multi-node serving architectures for distributed LLM inference.
- Optimize multi-LoRA serving systems.
- Apply advanced FP4 and FP6 quantization techniques to reduce model footprint while preserving quality.
- Implement speculative decoding and other latency optimization strategies.
- Develop disaggregated serving solutions with optimized caching for prefill and decoding phases.
- Benchmark and improve system performance across deployment scenarios and GPU types.
Requirements
- Extensive hands-on experience with state-of-the-art inference optimization techniques and scalable LLM inference systems.
- Proven expertise with distributed serving architectures for large language models.
- Hands-on experience implementing quantization techniques for transformer models, including FP4 and FP6.
- Strong understanding of speculative decoding with draft models and Eagle speculative decoding approaches.
- Proficiency in Python and C++.
- Experience with CUDA programming and GPU optimization.
- Preferred experience contributing to vLLM, SGLang, TensorRT-LLM, lmdeploy, or similar inference frameworks.
- Preferred experience developing custom CUDA kernels and deploying production inference systems.
- Deep understanding of performance optimization systems.
- Demonstrated experience building or optimizing LLM inference or training projects at scale.
Benefits
- Palo Alto office work five days per week to support collaboration and team culture.
- Opportunity to work on safety-focused healthcare LLM deployment at large scale.
- Opportunity to contribute to open-source inference frameworks and shape AI deployment technology.
Categories
About Hippocratic AI
Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.