ByteDance

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

ByteDance
Apply
2 hours ago
San Jose, CA, USAEntry Level
H1B sponsor

Responsibilities

  • Iterate on the architecture of a large-model inference engine and optimize GPU memory access, computation pipelines, asynchronous stream scheduling, throughput, and latency.
  • Adapt inference systems to GPU and NPU hardware architectures and improve hardware portability and performance.
  • Design, implement, and optimize distributed inference solutions using tensor, pipeline, sequence, and MoE expert parallelism.
  • Implement operator fusion, compilation optimization, memory reuse, scheduling optimization, quantization compilation, and other inference optimizations.
  • Benchmark and improve inference frameworks and develop solutions for high-concurrency, low-latency workloads.
  • Use GPU performance analysis tools to identify bottlenecks and drive software-hardware co-optimization and implementation iterations.
  • Collaborate across teams, document technical work, and resolve complex technical issues.

Requirements

  • Completing or recently completed a bachelor's or master's degree in software development, computer science, computer engineering, or a related technical discipline.
  • Strong low-level computer systems knowledge with proficiency in C/C++ and Python.
  • Proficiency in CUDA programming and familiarity with GPU architecture, memory models, computing scheduling, and communication mechanisms.
  • Ability to implement and optimize deep-learning operators, including matrix operations, normalization, and activation functions, with attention to memory access, vectorization, and precision alignment.
  • Familiarity with deep-learning inference compilation, including computational graph optimization, operator fusion, constant folding, memory reuse, scheduling optimization, and quantization compilation.
  • Proficiency with GPU performance analysis tools such as Nsight and Profiler.
  • Understanding of large-model inference, model parallelism, multi-card communication, load balancing, and parallel-efficiency optimization is preferred.
  • Experience with secondary development or performance optimization of vLLM, SGLang, TensorRT-LLM, or similar inference frameworks is preferred.
  • Strong collaboration, communication, presentation, documentation, responsibility, and problem-solving skills.

Benefits

  • Graduate role targeting a 2027 start.
  • Candidates must be able to commit to an onboarding date by the end of the year and should state availability and graduation date in their resume.
  • Applications are reviewed on a rolling basis, and candidates may apply to a maximum of two company or affiliate positions globally.

Tech Stack

ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me