ByteDance

Machine Learning Engineer Graduate (AML-Engine-Orchestration) - 2027 Start

ByteDance
Apply
2 hours ago

Responsibilities

  • Design and build ML platform orchestration capabilities using Kubernetes Operators, container runtimes, and lifecycle management for jobs, services, and stateful workloads.
  • Build multi-tenant resource and quota systems supporting priorities, preemption, fair sharing, elasticity, cross-cluster scheduling, resource pooling, and FinOps.
  • Improve GPU utilization and cost efficiency across heterogeneous compute infrastructure.
  • Develop online model-serving lifecycle orchestration for model and image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery.
  • Build serving orchestration and traffic-management capabilities including topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management.

Requirements

  • Completing or recently completed a bachelor’s or master’s degree in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field.
  • Proficiency in at least one of Go, C++, or Python, with a solid foundation in data structures, algorithms, and software engineering principles.
  • Familiarity with Linux and foundational knowledge of operating systems, computer networks, concurrent programming, and distributed systems.
  • Ability to investigate systems using source code, metrics, logs, profiling, and experiments.
  • Systematic and quantitative problem-solving skills, including defining measurements, testing hypotheses, and validating improvements.
  • Demonstrated ownership and collaboration through coursework, research, internships, open-source contributions, or other engineering projects.
  • Preferred experience with Kubernetes, container runtimes, resource scheduling, quota management, multi-tenant systems, or FinOps.
  • Preferred contributions to infrastructure projects such as Kubernetes, Volcano, Koordinator, or OpenKruise.
  • Preferred experience with vLLM, SGLang, Triton, KServe, Ray Serve, KV Cache, Continuous Batching, Prefill/Decode disaggregation, or model parallelism.
  • Preferred experience with online services, gateways, traffic management, autoscaling, performance optimization, highly available distributed systems, GPU/NPU programming, heterogeneous resource scheduling, model distribution, or inference performance analysis.

Benefits

  • Graduate opportunity with bold ideas, complex challenges, and growth opportunities.
  • Candidates must be able to commit to an onboarding date by the end of 2027 and should state availability and graduation date clearly in their resume.
  • Applications are reviewed on a rolling basis, and candidates may apply to a maximum of two company or affiliate positions globally.
ByteDance

About ByteDance

10,000+ employees

ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.

Contact me