Responsibilities
- Contribute to compiler optimizations for AI training and inference workloads.
- Develop and extend MLIR-based compiler passes for graph lowering, optimization, and code generation.
- Optimize model execution on GPU and NPU accelerators for performance, memory efficiency, and scalability.
- Support model deployment pipelines, including compilation, packaging, and runtime integration.
- Assist with distributed training and inference acceleration through parallel execution, communication optimization, and runtime scheduling.
- Benchmark, profile, and analyze large-scale model performance across hardware backends.
- Collaborate with researchers and engineers to translate model and system requirements into compiler and runtime improvements.
Requirements
- Currently pursuing a PhD in Computer Science, Electrical Engineering, or a related technical field.
- Experience using or developing open-source LLM inference frameworks such as vLLM or SGLang.
- Proficiency in at least one deep learning framework such as PyTorch, Megatron, DeepSpeed, or JAX, with experience in model inference workflows.
- Understanding of modern computing systems, including hardware, storage, networking, and their effects on machine-learning workloads.
- Familiarity with compilers, model optimization pipelines such as PyTorch Dynamo, or related model execution workflows.
- Experience with distributed or large-scale ML systems and training or inference optimizations such as FSDP, DeepSpeed, Megatron, or GSPMD is preferred.
- Experience with GPU, TPU, or NPU programming and performance optimization, or high-performance computing and communication using technologies such as CUDA, Triton, NCCL, or RDMA, is preferred.
- Understanding of AI compiler and model optimization stacks such as torch.fx, PyTorch Dynamo, XLA, or MLIR is preferred.
Benefits
- 12-week internship commitment in 2026.
- Hands-on learning, community-building and development events, and collaboration with industry experts.
- Applications are reviewed on a rolling basis, and applicants should state their availability and start and end dates in their resume.
Tech Stack
Categories
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
