2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Own production LLM inference services from model handoff through release engineering, capacity planning, cost optimization, and incident response.
- Tune inference performance to reduce end-to-end latency and increase throughput across real production traffic.
- Scale inference across heterogeneous GPU fleets and optimize runtimes, schedulers, KV caches, batching, and memory.
- Build benchmarking suites, metrics, and tooling for latency, throughput, GPU utilization, memory, and cost.
- Improve monitoring, tracing, alerting, incident response, and postmortem processes.
- Evaluate and implement inference optimizations such as quantization, paging, and kernel or runtime improvements.
- Partner with data science and product teams to translate business needs into performance and availability SLOs.
Requirements
- 5+ years of strong development experience.
- Experience deploying and operating LLM inference services in production.
- Strong production coding skills in Python plus Go or Rust, including systems-level implementation and debugging.
- Experience with PyTorch, vLLM, SGLang, and/or TensorRT.
- Knowledge of GPU architecture and performance, including profiling and memory bandwidth and latency tradeoffs; CUDA or kernel programming is a strong plus.
- Solid understanding of LLM inference and optimization techniques such as continuous batching, KV cache management, quantization, and speculative decoding.
- 3+ years of hands-on experience in performance optimization and systems programming for AI/ML workloads.
- Demonstrated ability to deliver measurable production improvements such as higher throughput, lower p95 or p99 latency, or reduced GPU cost.
- Proven root-cause analysis skills across models, runtimes, networking, and infrastructure.
