5 days ago
Responsibilities
- Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
- Optimize long-context prefill and decode workloads using real production traffic.
- Tune routing between internal infrastructure and external providers based on cost, capacity, and performance.
- Work with serving engines including vLLM, SGLang, and TensorRT-LLM, going below the framework when necessary.
- Build profiling and measurement systems to identify time, memory, and compute usage.
Requirements
- At least 5 years of experience in ML systems, inference infrastructure, or performance engineering, with measurable cost or latency improvements.
- Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
- Production experience with vLLM, SGLang, or TensorRT-LLM.
- Strong Python skills and proficiency in C++, Rust, or another systems language.
- Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
- Adaptability, collaborative teamwork, and willingness to develop bold ideas are valued.
Benefits
- Flexible work options including Bay Area in-person collaboration, a distributed global-first team, and team offsites.
- Annual Adaption Passport travel stipend for exploring a country not previously visited.
- Weekly lunch stipend for take-out or grocery delivery.
- Comprehensive medical benefits and generous paid time off.
Categories
About adaption
Adaption builds business-intelligence tools and guidance for climate-risk adaptation, helping companies and communities assess exposure and plan resilience measures. The privately held firm offers a software platform that translates environmental and infrastructure data into strategic insights, plus advisory support to turn findings into actionable plans. It serves organizations facing weather, regulatory, and supply-chain disruptions that need decision support for adaptation projects and investment planning.
