1 day ago
Singapore, SingaporeSenior
Responsibilities
- Profile, optimize, and extend LLM inference engines to improve throughput, latency, and GPU utilization.
- Implement inference optimizations including quantization, KV-cache management, continuous batching, speculative decoding, and tensor, expert, and pipeline parallelism.
- Develop and tune CUDA and Triton GPU kernels for attention, GEMM, and mixture-of-experts operators on modern NVIDIA architectures.
- Build benchmarking and load-testing frameworks for end-to-end serving performance under realistic workloads and SLOs.
- Deploy and scale inference services on Kubernetes-based GPU clusters with autoscaling, fault tolerance, and observability.
- Track inference research and bring advanced techniques into production.
Requirements
- At least 3 years of hands-on experience optimizing deep learning inference on NVIDIA GPUs, preferably for large language models.
- Strong knowledge of LLM inference internals, including attention mechanisms, KV caches, batching strategies, quantization, and parallelism.
- Experience with at least one inference serving engine such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or an equivalent.
- Proficiency in CUDA and/or Triton kernel programming and a solid understanding of modern GPU architecture.
- Strong software engineering fundamentals in Python and C++.
- Bachelor’s or higher degree in Computer Science, Engineering, or a related field.
- Experience contributing to open-source inference or ML systems projects is preferred.
- Familiarity with MoE serving, disaggregated prefill/decode, prefix caching, long-context optimization, low-precision inference, and model compression is preferred.
- Experience with Nsight Systems, Nsight Compute, PyTorch Profiler, Kubernetes, SLO-driven capacity planning, cost optimization, PyTorch, DeepSpeed, Megatron-LM, or FSDP is preferred.
Benefits
- Work on cutting-edge LLM inference optimization at scale.
- Directly impact Rakuten’s AI infrastructure by improving efficiency and reducing costs.
- Collaborate with global AI/ML teams on high-impact challenges.
- Research and implement state-of-the-art GPU optimizations.
Tech Stack
Categories
About Rakuten
Rakuten Group builds e-commerce marketplaces, financial services (credit cards, payments, banking, securities), and digital/communications offerings for consumers and merchants worldwide. It monetizes through transaction fees, financial services income, advertising, and mobile subscriptions; flagship businesses include Rakuten Ichiba, Rakuten Card, and Rakuten Mobile. Founded in 1997 and headquartered in Setagaya, Tokyo, the company is publicly listed on the Tokyo Stock Exchange and operates across many countries and regions.
