2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Build and operate production AI inference runtimes using vLLM and TensorRT-LLM behind a standardized AI runtime layer.
- Design and optimize distributed inference architectures, including prefill/decode disaggregation and distributed KV cache management.
- Develop large-scale distributed training pipelines using Megatron-LM and DeepSpeed on high-performance GPU clusters.
- Profile and resolve distributed training bottlenecks using NVIDIA and PyTorch performance tools.
- Implement inference optimizations including quantization, speculative decoding, continuous batching, and FlashAttention.
- Build and operate an inference request router for authentication, routing, and throughput management.
- Develop and operate multi-LoRA adapter hosting with hot-swap routing and lifecycle management.
- Build and maintain MLOps workflows for experiment tracking, model versioning, automated evaluation, and CI/CD.
- Develop and operate fine-tuning pipelines including SFT, RLHF, DPO, and LoRA.
- Build fault-tolerant distributed training infrastructure with checkpointing, failure detection, and recovery.
- Build regression testing and benchmarking systems to improve training and inference performance.
Requirements
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related field.
- 5+ years of experience building distributed systems or ML platform infrastructure.
- Strong programming skills in Python and/or Golang.
- Hands-on experience deploying and operating LLM inference engines such as vLLM, TensorRT-LLM, NVIDIA Triton, or SGLang.
- Deep understanding of LLM inference internals, including KV cache management, PagedAttention, continuous batching, and request routing.
- Experience building or optimizing distributed training pipelines using Megatron-LM, DeepSpeed, FSDP, or equivalent frameworks.
- Strong understanding of model parallelism strategies and their trade-offs.
- Proficiency with NVIDIA tooling such as NCCL, DCGM, Nsight Systems, and PyTorch Profiler.
- Experience implementing inference optimizations including quantization, speculative decoding, FlashAttention, and multi-LoRA serving.
- Experience building MLOps workflows including experiment tracking, model registry, evaluation automation, and CI/CD.
- Experience developing fine-tuning pipelines such as SFT, RLHF, DPO, or LoRA at scale.
- Strong expertise in Kubernetes and containerized GPU environments.
- Strong debugging and performance optimization skills across CUDA runtimes, distributed training, and ML serving systems.
- Familiarity with CUDA or Triton kernel development is a plus.
