Responsibilities
- Design and build machine learning platform orchestration capabilities using Kubernetes Operators, container runtimes, and lifecycle management for jobs, services, and stateful workloads.
- Build multi-tenant resource and quota systems supporting priorities, preemption, fair sharing, elasticity, and cross-cluster scheduling.
- Improve GPU utilization and cost efficiency through resource pooling and FinOps.
- Develop online model-serving lifecycle orchestration covering model and image distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery.
- Build serving orchestration and traffic-management capabilities for disaggregated serving clusters, including topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management.
Requirements
- Be completing or have recently completed a bachelor's or master's degree in computer science, software engineering, artificial intelligence, or a related technical field.
- Be proficient in at least one of Go, C++, or Python and have a solid foundation in data structures, algorithms, and software engineering principles.
- Have familiarity with Linux and foundational knowledge of operating systems, computer networks, concurrent programming, and distributed systems.
- Demonstrate hands-on exploratory ability using source code, metrics, logs, profiling, and experiments.
- Use a systematic and quantitative approach to define measurements, test hypotheses, and validate system improvements.
- Demonstrate ownership and collaboration through coursework, research, internships, open-source contributions, or other engineering projects.
- Preferred qualifications include experience with Kubernetes, container runtimes, resource scheduling, quota management, multi-tenant systems, or FinOps.
- Preferred qualifications include contributions to infrastructure projects such as Kubernetes, Volcano, Koordinator, or OpenKruise.
- Preferred qualifications include experience with vLLM, SGLang, Triton, KServe, Ray Serve, or concepts such as KV Cache, Continuous Batching, Prefill/Decode disaggregation, and model parallelism.
- Preferred qualifications include experience with online services, gateways, traffic management, autoscaling, performance optimization, or highly available distributed systems.
- Preferred qualifications include experience with GPU/NPU programming, heterogeneous resource scheduling, model distribution, or inference performance analysis.
Benefits
- Graduate role with a 2027 start and an onboarding date commitment required by the end of the year.
- Applications are reviewed on a rolling basis, and candidates may apply to a maximum of two positions globally across the company and its affiliates.
Tech Stack
Categories
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
