Responsibilities
- Design and build ML platform orchestration capabilities for jobs, services, and stateful workloads.
- Develop multi-tenant resource, quota, priority, preemption, fair-sharing, elasticity, and cross-cluster scheduling systems.
- Improve GPU utilization and cost efficiency through resource pooling and FinOps.
- Build model-serving lifecycle capabilities including distribution, deployment, upgrades, rollback, autoscaling, multi-cluster operation, and disaster recovery.
- Develop topology-aware scheduling, KV Cache affinity, intelligent request routing, and QoS/SLA management for disaggregated serving clusters.
Requirements
- Completing or recently completed a PhD in Computer Science, Software Engineering, Artificial Intelligence, or a related technical field.
- Proficiency in at least one of Go, C++, or Python, with a foundation in data structures, algorithms, and software engineering principles.
- Familiarity with Linux, operating systems, computer networks, concurrent programming, and distributed systems.
- Hands-on exploratory ability using source code, metrics, logs, profiling, and experiments.
- A systematic and quantitative approach to problem solving, including defining measurements, testing hypotheses, and validating improvements.
- Demonstrated ownership and collaboration through coursework, research, internships, open-source contributions, or engineering projects.
- Preferred experience with Kubernetes, container runtimes, resource scheduling, quota management, multi-tenant systems, or FinOps.
- Preferred contributions to infrastructure projects such as Kubernetes, Volcano, Koordinator, or OpenKruise.
- Preferred experience with vLLM, SGLang, Triton, KServe, Ray Serve, KV Cache, Continuous Batching, Prefill/Decode disaggregation, or model parallelism.
- Preferred experience with online services, gateways, traffic management, autoscaling, performance optimization, highly available distributed systems, GPU/NPU programming, heterogeneous resource scheduling, model distribution, or inference performance analysis.
Benefits
- Graduate opportunity with an onboarding date that must be committed to by the end of the year; candidates should state availability and graduation date in their resume.
Tech Stack
Categories
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
