Responsibilities
- Design and evaluate scalable AI factory architectures across compute, storage, networking, chips, power, data, and application layers.
- Develop technical proposals addressing supply-chain and energy constraints alongside silicon and software trade-offs.
- Track emerging trends in AI systems, distributed training, reinforcement learning, hardware acceleration, cognitive science, and psychology.
- Build prototypes and communicate findings through technical reports.
- Benchmark and optimize scheduling, networking, storage, training and RL frameworks, and AI memory systems.
- Collaborate with research, engineering, hardware, data-center, and product teams to drive cross-team AI infrastructure initiatives.
Requirements
- Completing or recently completed a PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical discipline.
- Candidates from cognitive science, computational neuroscience, or psychology should also have strong systems fundamentals.
- Experience with distributed systems, infrastructure engineering, or ML systems, including exposure to large-scale training or RL pipelines.
- Ability to evaluate trade-offs across hardware, software, algorithms, energy, and supply-chain constraints.
- Strong proficiency integrating AI tools into knowledge discovery and research workflows.
- Ability to learn quickly and remain productive in a rapidly evolving technical area.
- Strong communication and cross-functional collaboration skills.
- Preferred experience with large-scale model training and inference, distributed pretraining, post-training, RL, KV-cache-aware serving, GPU or accelerator optimization, and high-performance networking such as RDMA and NCCL.
- Preferred experience with heterogeneous AI compute systems, large-scale training clusters, HPC-style distributed workloads, and training and evaluation data pipelines.
- Familiarity with AI memory systems, retrieval-augmented architectures, or long-term agent memory designs.
- Exposure to chip-level design, data-center energy and cooling, or AI hardware supply-chain considerations.
- Publications in systems or machine learning conferences and contributions to open-source projects are preferred.
Categories
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
