Hippocratic AI

LLM Inference Systems Engineer

Hippocratic AI
Apply
14 hours ago
Menlo Park, CA, USASenior
H1B sponsor

Responsibilities

  • Design, implement, and operate disaggregated LLM serving architectures with prefill/decode pools, KV-cache transfer, routing, batching, and failure recovery.
  • Optimize multi-node inference across GPU communication and networking systems.
  • Build and tune serving systems with SGLang, vLLM, NVIDIA Dynamo, TensorRT-LLM, or comparable frameworks.
  • Profile and improve time to first token, inter-token latency, throughput, tail latency, and cost per token.
  • Develop benchmark and capacity-planning workflows across models, GPU types, parallelism strategies, and concurrency levels.
  • Productionize serving on Kubernetes with health checks, autoscaling, safe rollouts, metrics, tracing, and diagnostics.
  • Diagnose distributed failures including collective timeouts, topology mismatches, packet loss, congestion, KV-transfer stalls, GPU out-of-memory errors, and uneven load.
  • Partner with model researchers and platform engineers to launch models, serving features, and hardware generations safely.

Requirements

  • Production experience building or operating distributed systems, high-performance computing systems, or large-scale ML inference platforms.
  • Strong understanding of LLM inference, including tensor, pipeline, and data parallelism, continuous batching, KV-cache management, and prefill/decode behavior.
  • Hands-on experience with GPU communication and networking technologies such as NCCL, NVLink/NVSwitch, InfiniBand, RoCE, RDMA, or UCX.
  • Strong Python skills and working proficiency in C++ or another systems language.
  • Ability to profile and debug performance across application, runtime, kernel, network, and infrastructure layers.
  • Experience deploying and operating production workloads on Kubernetes with reliability and observability standards.
  • Clear written and verbal communication across research, infrastructure, and product-facing teams.
  • Preferred experience with disaggregated serving, KV-cache transfer, NVIDIA Dynamo/NIXL, SGLang, vLLM, or TensorRT-LLM.
  • Preferred experience with NVLS, SHARP, GPUDirect RDMA, UCX, RDMA congestion control, or GPU-cluster topology optimization.
  • Preferred CUDA, Triton, custom-kernel, or low-level GPU performance experience.
  • Preferred experience with speculative decoding, multi-LoRA serving, quantization, cache-aware routing, open-source inference or networking projects, and benchmarking new GPU platforms.

Benefits

  • The role is based in the Menlo Park office five days per week unless otherwise specified.
  • Employees will work on a safety-focused healthcare AI platform alongside physicians, researchers, and engineering experts.
  • Hippocratic AI is an equal opportunity employer and provides accommodations during the hiring process.
Hippocratic AI

About Hippocratic AI

201-500 employees

Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.

Contact me