Responsibilities
- Set technical direction and own a multi-quarter roadmap across post-training, agent harnesses, evaluation, self-evolving data systems, or serving.
- Translate north-star metrics into 2–3 high-ROI technical bets per quarter and deliver them.
- Hire and develop 1–3 strong individual contributors and raise the technical bar through design reviews, code reviews, and mentoring.
- Write load-bearing production code, define core abstractions, and author design documents.
- Drive alignment with foundation-model, infrastructure, product, and adjacent algorithm teams and own sign-off on cross-cutting technical decisions.
- Build per-turn tracing, tool-call analytics, failure-mode taxonomies, observability, and rollback systems.
Requirements
- BS, MS, or PhD in Computer Science, AI, Mathematics, or a related quantitative field.
- Hands-on experience in machine learning, natural language processing, or applied deep learning; top PhDs with strong publication records may qualify with 4+ years of experience.
- Strong Python and production-path experience with at least one of C++, Go, or Rust.
- Hands-on post-training or fine-tuning of frontier-class LLMs at 7B-plus scale across multiple nodes; API-only experience is insufficient.
- Experience leading at least one production LLM or agent system from zero to one.
- Preferred experience includes multi-node SFT, DPO, online RL, reward modeling, preference data construction, RLAIF, RLVR, distillation, QAT, long-context training, and continual pre-training.
- Preferred experience includes production agent harnesses, context engineering, sub-agents, durable execution, MCP or Skill-style extensibility, parallel tool use, and computer use.
- Preferred experience includes reasoning-trace training, planner/critic decomposition, self-consistency, verifier models, and multi-step reasoning evaluation.
- Preferred experience includes LLM-as-judge evaluation, human-agreement calibration, benchmark harnesses, regression, safety, cost, and latency evaluation.
- Preferred experience includes vLLM or TensorRT-LLM, mixture-of-experts serving, speculative decoding, KV-cache and prompt caching, and low-latency multi-tenant serving.
- Preferred experience includes multilingual SFT/DPO, low-resource adaptation, faithful machine translation, and locale-aware reasoning.
Categories
About ByteDance
ByteDance is a global incubator of platforms at the cutting edge of commerce, content, entertainment and enterprise services - over 2.5bn people interact with ByteDance products including TikTok. Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. We are committed to building a safe, healthy and positive online environment for all our users. We have over 110,000 employees based in more than 30 countries globally. Join us.
