ByteDance

Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)

ByteDance
Apply
4 hours ago
Singapore, SingaporeMid Level

Responsibilities

  • Iterate on the architecture of the large-model inference engine and optimize GPU performance, memory access, computation pipelines, and asynchronous scheduling.
  • Adapt the inference engine to GPU and NPU hardware architectures and improve throughput, latency, and hardware utilization.
  • Design, develop, and optimize distributed inference solutions using tensor, pipeline, sequence, and MoE expert parallelism.
  • Address multi-card model splitting, cross-card communication overhead, load imbalance, and parallel efficiency issues.
  • Research and implement advances in large-model inference, GPU high-performance computing, distributed parallelism, cache optimization, and inference frameworks.
  • Benchmark and optimize inference systems to improve performance and cost efficiency for high-concurrency, low-latency workloads.

Requirements

  • Bachelor’s degree in Computer Science or equivalent with 3+ years of relevant experience.
  • Strong low-level computer systems knowledge and proficiency in C/C++, Python, and CUDA, with familiarity with GPU architecture, memory models, computation scheduling, and communication mechanisms.
  • Experience implementing and optimizing deep-learning operators, including matrix operations, normalization, and activation functions, through memory optimization, vectorization, handwritten reconstruction, and precision alignment.
  • Understanding of inference compilation technologies including computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation.
  • Proficiency with GPU performance analysis tools such as Nsight and Profiler and the ability to identify and resolve inference bottlenecks.
  • Understanding of large-model inference, model parallelism, distributed inference, multi-card communication, load balancing, and parallel-efficiency optimization is preferred.
  • Experience with secondary development and performance optimization of vLLM, SGLang, or TensorRT-LLM is preferred.
  • Familiarity with computation-efficiency optimization for mainstream deep-learning frameworks is preferred.

Tech Stack

ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me