TCGplayer

AI Platform Engineer

TCGplayer
Apply
2 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Build and operate production AI inference runtimes using vLLM and TensorRT-LLM behind a standardized AI runtime layer.
  • Design and optimize distributed inference architectures, including prefill/decode disaggregation and distributed KV cache management.
  • Develop large-scale distributed training pipelines using Megatron-LM and DeepSpeed on high-performance GPU clusters.
  • Profile and resolve distributed training bottlenecks using NVIDIA and PyTorch performance tools.
  • Implement inference optimizations including quantization, speculative decoding, continuous batching, and FlashAttention.
  • Build and operate an inference request router for authentication, routing, and throughput management.
  • Develop and operate multi-LoRA adapter hosting with hot-swap routing and lifecycle management.
  • Build and maintain MLOps workflows for experiment tracking, model versioning, automated evaluation, and CI/CD.
  • Develop and operate fine-tuning pipelines including SFT, RLHF, DPO, and LoRA.
  • Build fault-tolerant distributed training infrastructure with checkpointing, failure detection, and recovery.
  • Build regression testing and benchmarking systems to improve training and inference performance.

Requirements

  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience building distributed systems or ML platform infrastructure.
  • Strong programming skills in Python and/or Golang.
  • Hands-on experience deploying and operating LLM inference engines such as vLLM, TensorRT-LLM, NVIDIA Triton, or SGLang.
  • Deep understanding of LLM inference internals, including KV cache management, PagedAttention, continuous batching, and request routing.
  • Experience building or optimizing distributed training pipelines using Megatron-LM, DeepSpeed, FSDP, or equivalent frameworks.
  • Strong understanding of model parallelism strategies and their trade-offs.
  • Proficiency with NVIDIA tooling such as NCCL, DCGM, Nsight Systems, and PyTorch Profiler.
  • Experience implementing inference optimizations including quantization, speculative decoding, FlashAttention, and multi-LoRA serving.
  • Experience building MLOps workflows including experiment tracking, model registry, evaluation automation, and CI/CD.
  • Experience developing fine-tuning pipelines such as SFT, RLHF, DPO, or LoRA at scale.
  • Strong expertise in Kubernetes and containerized GPU environments.
  • Strong debugging and performance optimization skills across CUDA runtimes, distributed training, and ML serving systems.
  • Familiarity with CUDA or Triton kernel development is a plus.
TCGplayer

About TCGplayer

201-500 employees
Contact me