TCGplayer

AI Platform Engineer

TCGplayer
Apply
2 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Own production LLM inference services from model handoff through release engineering, capacity planning, cost optimization, and incident response.
  • Tune inference performance to reduce end-to-end latency and increase throughput across real production traffic.
  • Scale inference across heterogeneous GPU fleets and optimize runtimes, schedulers, KV caches, batching, and memory.
  • Build benchmarking suites, metrics, and tooling for latency, throughput, GPU utilization, memory, and cost.
  • Improve monitoring, tracing, alerting, incident response, and postmortem processes.
  • Evaluate and implement inference optimizations such as quantization, paging, and kernel or runtime improvements.
  • Partner with data science and product teams to translate business needs into performance and availability SLOs.

Requirements

  • 5+ years of strong development experience.
  • Experience deploying and operating LLM inference services in production.
  • Strong production coding skills in Python plus Go or Rust, including systems-level implementation and debugging.
  • Experience with PyTorch, vLLM, SGLang, and/or TensorRT.
  • Knowledge of GPU architecture and performance, including profiling and memory bandwidth and latency tradeoffs; CUDA or kernel programming is a strong plus.
  • Solid understanding of LLM inference and optimization techniques such as continuous batching, KV cache management, quantization, and speculative decoding.
  • 3+ years of hands-on experience in performance optimization and systems programming for AI/ML workloads.
  • Demonstrated ability to deliver measurable production improvements such as higher throughput, lower p95 or p99 latency, or reduced GPU cost.
  • Proven root-cause analysis skills across models, runtimes, networking, and infrastructure.
TCGplayer

About TCGplayer

201-500 employees
Contact me