SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading
Xingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang, Shuai Wang, Shangguang Wang, Zhehao Lin, Yinben Xia, Chang Yu, Qihang Liu, Xuan Zhang, Hao Lu
Abstract
Large Language Models (LLMs) increasingly rely on Mixture-of-Experts (MoE) architectures to scale computation efficiently. Expert Parallelism (EP), which distributes experts across GPUs, introduces all-to-all communication overhead during the dispatch and combine phases, especially in the prefill stage, which dominates the inference performance. Existing communication libraries, such as DeepEP, suffer from excessive GPU SM utilization and underutilized interconnect bandwidth, limiting prefill performance.
In this paper, we identify two root causes: redundant buffer copies and inefficient intra-server transfers over NVLink. To address these, we propose SwiftEP, an all-to-all communication library tailored for MoE prefill, combining buffer fusion and Tensor Memory Accelerator (TMA) offloading. Buffer fusion eliminates redundant staging copies, enabling true zero-copy communication, while TMA offloading maximizes NVLink utilization and supports efficient multicast/reduce operations. SwiftEP further incorporates RDMA scatter-gather lists, QP transmission parallelization, and CUDA IPC to handle dynamic token placement and inter-GPU memory access. Evaluation on 16-and 32-GPU clusters shows that SwiftEP achieves up to 119.7% higher algorithm bandwidth, reduces SM occupancy by up to 66.7%, and improves request serving capacity by 21.2% compared to DeepEP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7219e925-fa80-4a06-9009-d6fd3b925a1bBuilds on14
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement LearningHung Le, Yue Wang, Akhilesh Deepak Gotmare, Silvio Savarese et al.NeurIPS 2022 · 571 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
Related papers
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu et al.SIGCOMM 2025 · 19 citations
- UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsYipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng et al.SIGCOMM 2026
- Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM InferenceZhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou et al.ISCA 2026
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin et al.NeurIPS 2025 · 6 citations
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-SchedulingYan Li, Zhenyu Zhang, Zhengang Wang, Pengfei chen et al.ICLR 2026 · 11 citations
