
LLM Inference Engineer
Hippocratic AIResponsibilities
- Design and implement multi-node serving architectures for distributed LLM inference.
- Optimize multi-LoRA serving systems.
- Apply advanced FP4 and FP6 quantization techniques to reduce model footprint while preserving quality.
- Implement speculative decoding and other latency optimization strategies.
- Develop disaggregated serving solutions with optimized caching for prefill and decoding phases.
- Benchmark and improve system performance across deployment scenarios and GPU types.
Requirements
- Extensive hands-on experience with state-of-the-art inference optimization techniques and scalable LLM inference systems.
- Proven expertise with distributed serving architectures for large language models.
- Hands-on experience implementing quantization techniques for transformer models, including FP4 and FP6.
- Strong understanding of speculative decoding with draft models and Eagle speculative decoding approaches.
- Proficiency in Python and C++.
- Experience with CUDA programming and GPU optimization.
- Preferred experience contributing to vLLM, SGLang, TensorRT-LLM, lmdeploy, or similar inference frameworks.
- Preferred experience developing custom CUDA kernels and deploying production inference systems.
- Deep understanding of performance optimization systems.
- Demonstrated experience building or optimizing LLM inference or training projects at scale.
Benefits
- Palo Alto office work five days per week to support collaboration and team culture.
- Opportunity to work on safety-focused healthcare LLM deployment at large scale.
- Opportunity to contribute to open-source inference frameworks and shape AI deployment technology.
Categories
About Hippocratic AI
Hippocratic AI has developed a safety-focused Large Language Model (LLM) for healthcare. The company believes that a safe LLM can dramatically improve healthcare accessibility and health outcomes in the world by bringing deep healthcare expertise to every human. No other technology has the potential to have this level of global impact on health. The company was co-founded by CEO Munjal Shah, alongside a group of physicians, hospital administrators, healthcare professionals, and artificial intelligence researchers from El Camino Health, Johns Hopkins, Stanford, Microsoft, Google, and NVIDIA. Hippocratic AI has received a total of $278 million in funding and is backed by leading investors, including Andreessen Horowitz, General Catalyst, Kleiner Perkins, NVIDIA’s NVentures, Premji Invest, SV Angel, and six health systems. For more information on Hippocratic AI, www.HippocraticAI.com. Be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from @hippocraticai.com email addresses. We will never request payment or sensitive personal information. If anything appears suspicious, stop engaging immediately and report the incident.