Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM Inference
Zhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou, Zhengding Hu, Shuyi Pei, Yangwook Kang, Yufei Ding, Po-An Tsai
Abstract
Large-scale Mixture of Experts (MoE) Large Language Models (LLMs) have recently become the frontier open-weight models, achieving remarkable model capability similar to proprietary ones. But their random expert selection mechanism introduces significant data movement overhead that becomes the dominant bottleneck in multi-unit LLM serving systems. To understand the patterns underlying this data movement, we conduct comprehensive data-movement-centric profiling across four state-of-the-art large-scale MoE models released in 2025 (200B-1000B) using over 24,000 requests spanning diverse workloads. We perform systematic analysis from both temporal and spatial perspectives and distill six key insights to guide the design of diverse serving systems. We verify these insights on both future wafer-scale GPU architectures and existing GPU systems. On wafer-scale GPUs, lightweight architectural modifications guided by our insights yield a 6.6 average speedup across four 200B--1000B models. On existing GPU systems, our insights drive the design of a prefill-aware expert placement algorithm that achieves up to 1.25 speedup on MoE computation. Our work presents the first comprehensive data-centric analysis of large-scale MoE models together with a concrete design study applying the learned lessons. Our profiling traces are publicly available at bluehttps://huggingface.co/datasets/core12345/MoE_expert_selection_trace.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afb1eddc-b827-4a5c-b0c2-8601a22da805Builds on38
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 1,224 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
Related papers
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-SchedulingYan Li, Zhenyu Zhang, Zhengang Wang, Pengfei chen et al.ICLR 2026 · 11 citations
- Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert CachingKexin Li, Wenkan Huang, Qinggang Wang, Long Zheng et al.SC 2025 · 3 citations
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang et al.ISCA 2026
- MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource ConstraintsYichao Yuan, Lin Ma, Nishil TalatiHPDC 2026
