Hippocratic AI

LLM Inference Engineer

Hippocratic AI
Apply
11 months ago
Palo Alto, CA, USASenior
H1B sponsor

Responsibilities

  • Design and implement multi-node serving architectures for distributed LLM inference.
  • Optimize multi-LoRA serving systems.
  • Apply advanced FP4 and FP6 quantization techniques to reduce model footprint while preserving quality.
  • Implement speculative decoding and other latency optimization strategies.
  • Develop disaggregated serving solutions with optimized caching for prefill and decoding phases.
  • Benchmark and improve system performance across deployment scenarios and GPU types.

Requirements

  • Extensive hands-on experience with state-of-the-art inference optimization techniques and scalable LLM inference systems.
  • Proven expertise with distributed serving architectures for large language models.
  • Hands-on experience implementing quantization techniques for transformer models, including FP4 and FP6.
  • Strong understanding of speculative decoding with draft models and Eagle speculative decoding approaches.
  • Proficiency in Python and C++.
  • Experience with CUDA programming and GPU optimization.
  • Preferred experience contributing to vLLM, SGLang, TensorRT-LLM, lmdeploy, or similar inference frameworks.
  • Preferred experience developing custom CUDA kernels and deploying production inference systems.
  • Deep understanding of performance optimization systems.
  • Demonstrated experience building or optimizing LLM inference or training projects at scale.

Benefits

  • Palo Alto office work five days per week to support collaboration and team culture.
  • Opportunity to work on safety-focused healthcare LLM deployment at large scale.
  • Opportunity to contribute to open-source inference frameworks and shape AI deployment technology.

Tech Stack

Hippocratic AI

About Hippocratic AI

201-500 employees

Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.

Contact me