
LLM Inference Systems Engineer
Hippocratic AI14 hours ago
Responsibilities
- Design, implement, and operate disaggregated LLM serving architectures with prefill/decode pools, KV-cache transfer, routing, batching, and failure recovery.
- Optimize multi-node inference across GPU communication and networking systems.
- Build and tune serving systems with SGLang, vLLM, NVIDIA Dynamo, TensorRT-LLM, or comparable frameworks.
- Profile and improve time to first token, inter-token latency, throughput, tail latency, and cost per token.
- Develop benchmark and capacity-planning workflows across models, GPU types, parallelism strategies, and concurrency levels.
- Productionize serving on Kubernetes with health checks, autoscaling, safe rollouts, metrics, tracing, and diagnostics.
- Diagnose distributed failures including collective timeouts, topology mismatches, packet loss, congestion, KV-transfer stalls, GPU out-of-memory errors, and uneven load.
- Partner with model researchers and platform engineers to launch models, serving features, and hardware generations safely.
Requirements
- Production experience building or operating distributed systems, high-performance computing systems, or large-scale ML inference platforms.
- Strong understanding of LLM inference, including tensor, pipeline, and data parallelism, continuous batching, KV-cache management, and prefill/decode behavior.
- Hands-on experience with GPU communication and networking technologies such as NCCL, NVLink/NVSwitch, InfiniBand, RoCE, RDMA, or UCX.
- Strong Python skills and working proficiency in C++ or another systems language.
- Ability to profile and debug performance across application, runtime, kernel, network, and infrastructure layers.
- Experience deploying and operating production workloads on Kubernetes with reliability and observability standards.
- Clear written and verbal communication across research, infrastructure, and product-facing teams.
- Preferred experience with disaggregated serving, KV-cache transfer, NVIDIA Dynamo/NIXL, SGLang, vLLM, or TensorRT-LLM.
- Preferred experience with NVLS, SHARP, GPUDirect RDMA, UCX, RDMA congestion control, or GPU-cluster topology optimization.
- Preferred CUDA, Triton, custom-kernel, or low-level GPU performance experience.
- Preferred experience with speculative decoding, multi-LoRA serving, quantization, cache-aware routing, open-source inference or networking projects, and benchmarking new GPU platforms.
Benefits
- The role is based in the Menlo Park office five days per week unless otherwise specified.
- Employees will work on a safety-focused healthcare AI platform alongside physicians, researchers, and engineering experts.
- Hippocratic AI is an equal opportunity employer and provides accommodations during the hiring process.
Tech Stack
Categories
About Hippocratic AI
Hippocratic AI builds a safety-focused large language model and AI agents for healthcare workflows, used by health systems for patient outreach, post-discharge follow-up, and chronic-care management. It licenses its platform and tools to providers to automate and scale clinical support tasks while meeting health-system requirements. Founded in 2023 and headquartered in Palo Alto, the company is privately held.