StreamEP: Straggler-Tolerant MoE Decoding without Communication Barriers
Yizhuo Liang, Shaoyu Wang, Jaeyong Song, Yanqi Zhou, Geon-Woo Kim, Guangrong He, Seo Jin Park
Abstract
Mixture-of-Experts (MoE) models use expert parallelism (EP) to distribute experts across GPUs, relying on collective communication that synchronizes every device at every layer. To minimize communication latency, hyperscalers provision abundant network bandwidth in data centers and deploy MoE-optimized stacks that require Hopper-generation GPU features. But on commodity GPU clusters with shared and nonuniform interconnects, this communication barrier serializes execution behind the slowest transfer.
We present stream expert parallelism (StreamEP), a paradigm that removes the collective communication barrier entirely: each GPU independently receives tokens, executes the appropriate layer, and forwards results as they become ready. We build StreamInfer, a serving system that realizes StreamEP on vanilla NCCL. StreamInfer introduces per-layer token queues for asynchronous token management, a defragmenting scheduler that consolidates partial batches while guaranteeing forward progress, and a two-phase communicator for asynchronous and variable-size GPU-to-GPU transfers. On 16×L40S and 16×A100 clusters with 200 Gbps internode networking, StreamInfer achieves up to 1.7× higher decoding throughput than traditional EP, demonstrating that efficient large-scale MoE serving is achievable on commodity hardware.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8023b481-ce68-46f9-bdd2-b6e989a1067aBuilds on16
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-SchedulingYan Li, Zhenyu Zhang, Zhengang Wang, Pengfei chen et al.ICLR 2026 · 11 citations
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu et al.SIGCOMM 2025 · 19 citations
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 22 citations
- Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU–GPU Hybrid DesignWenxin Wang, Yule Hou, Yu Ji, Peng Qu et al.OSDI 2026 · 1 citation
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
