ByteDance

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance
Apply
18 hours ago
Singapore, SingaporeEntry Level

Responsibilities

  • Iterate on the underlying architecture of large-model inference engines and optimize end-to-end GPU performance.
  • Optimize GPU memory access, computing pipelines, asynchronous Stream scheduling, inference throughput, and latency.
  • Adapt inference engines to GPU and NPU hardware architectures and improve hardware universality.
  • Design and implement distributed inference parallelism using tensor, pipeline, sequence, and MoE expert parallelism.
  • Research and implement improvements involving large-model inference, GPU high-performance computing, distributed parallelism, cache optimization, and mainstream inference frameworks.
  • Benchmark and optimize inference-system performance, efficiency, and cost.

Requirements

  • Be completing or have recently completed a bachelor’s or master’s degree in computing or a related discipline.
  • Have strong low-level computer systems knowledge and proficiency in C/C++ and Python.
  • Be skilled in CUDA programming and familiar with GPU architecture, memory models, computing scheduling, and communication mechanisms.
  • Understand deep-learning operators and GPU optimization techniques including memory-access optimization, vectorization, handwritten operator reconstruction, and precision alignment.
  • Understand inference compilation technologies including computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation.
  • Be able to use GPU performance analysis tools such as Nsight and Profiler to identify and resolve inference bottlenecks.
  • Demonstrate collaboration, communication, presentation, documentation, responsibility, and technical problem-solving skills.
  • Preferred: experience with model parallelism, distributed inference, multi-card communication, load balancing, and parallel-efficiency optimization.
  • Preferred: secondary-development or performance-optimization experience with vLLM, SGLang, or TensorRT-LLM.

Benefits

  • Graduate opportunity with a planned 2027 start.
  • Candidates must be able to commit to an onboarding date by the end of 2027.
  • Applications are reviewed on a rolling basis, and candidates should state their availability and graduation date in their resume.
  • Candidates may apply to a maximum of two ByteDance or affiliate positions globally.

Tech Stack

ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me