4 hours ago
Responsibilities
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads using real production traffic.
- Tune routing between internal infrastructure and external providers based on cost, capacity, and performance.
- Work with vLLM, SGLang, and TensorRT-LLM serving engines, including below-framework optimization when needed.
- Build profiling and measurement systems to identify time, memory, and compute usage.
Requirements
- 5+ years of experience in ML systems, inference infrastructure, or performance engineering, with measurable cost or latency improvements.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
- Adaptability, collaborative ability, and willingness to contribute bold ideas are valued.
Benefits
- Flexible work options include in-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
- Annual Adaption Passport travel stipend to explore a country never visited before.
- Weekly lunch stipend for take-out or grocery delivery.
- Comprehensive medical benefits and generous paid time off.
Categories
About adaption
Adaption builds business-intelligence tools and guidance for climate-risk adaptation, helping companies and communities assess exposure and plan resilience measures. The privately held firm offers a software platform that translates environmental and infrastructure data into strategic insights, plus advisory support to turn findings into actionable plans. It serves organizations facing weather, regulatory, and supply-chain disruptions that need decision support for adaptation projects and investment planning.
