11 hours ago
Shanghai, ChinaStaff+
Responsibilities
- Lead the technical direction and continuous evolution of a discrete-event simulation platform for distributed multi-accelerator inference systems.
- Model continuous batching, chunked prefill, prefill-decode disaggregation, dynamic KV cache allocation, interconnect behavior, and collective communications.
- Design automated pipelines for empirical performance data ingestion, parameter sweeps, and results analysis.
- Drive production trace replay and sensitivity analysis to identify compute, memory, interconnect, and scheduling bottlenecks.
- Evaluate inference SLOs, throughput, power, and cost trade-offs across realistic and agentic workloads.
- Translate simulation findings into AI SoC and platform architecture recommendations for compute, memory, CPU offloading, and interfaces.
- Mentor engineers on simulation methods, code quality, calibration, and experimental rigor, while coordinating with silicon, software, and platform teams.
Requirements
- Ph.D. or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related quantitative field; a Bachelor's degree is listed with a higher experience requirement.
- Ph.D. candidates need 4–6+ years, Master's candidates 6–8+ years, or Bachelor's candidates 8–10+ years in system performance modeling, computer architecture, or AI systems engineering.
- Hands-on experience with discrete-event, trace-driven, or analytical performance simulators is required.
- Demonstrated ability to identify performance bottlenecks across compute, memory, and interconnects in multi-accelerator or heterogeneous environments.
- Proficiency in Python and C++14/17/20, with experience in modular software design and automated data processing.
- Understanding of LLM serving runtimes such as vLLM, TensorRT-LLM, and SGLang, plus mechanisms including PagedAttention, continuous batching, and speculative decoding is preferred.
- Knowledge of UALink, PCIe Gen 6/7, CXL, RoCEv2, or Ultra Ethernet is preferred.
- Experience with host-accelerator partitioning, CPU-assisted pipelines, tiered memory, production trace telemetry, MoE routing, or agentic workflows is preferred.
- A track record of leading technical projects, driving team consensus, and mentoring engineers is preferred.
Benefits
- Hybrid work model allowing time on-site at the assigned Intel site and off-site.
- Shift 1 position based in Shanghai, China.
About Intel
Intel is a public semiconductor company that designs and manufactures CPUs and other silicon for PCs, servers, networking, edge, and embedded systems, and also offers foundry services plus AI and graphics accelerators. Founded in 1968 and headquartered in Santa Clara, it sells primarily to OEMs, cloud providers, and device makers worldwide. Its portfolio includes x86 processors (Core and Xeon), Ethernet and chipset products, and the Mobileye autonomous driving platform.
