FlashMoE: Fast Distributed MoE in a Single Kernel
Osayamen Jonathan Aimuyo, Byungsoo Oh, Rachee Singh
Abstract
The computational sparsity of Mixture-of-Experts (MoE) models enables sub-linear growth in compute cost as model size increases, thus offering a scalable path to training massive neural networks. However, existing implementations suffer from low GPU utilization, significant latency overhead, and a fundamental inability to leverage task locality, primarily due to CPU-managed scheduling, host-initiated communication, and frequent kernel launches. To overcome these limitations, we develop FlashMoE, a fully GPU-resident MoE operator that fuses expert computation and inter-GPU communication into a single persistent GPU kernel. FlashMoE enables fine-grained pipelining of dispatch, compute, and combine phases, eliminating launch overheads and reducing idle gaps. Unlike existing work, FlashMoE eliminates bulk-synchronous collectives for one-sided, device-initiated, inter-GPU (R)DMA transfers, thereby unlocking payload efficiency by eliminating bloated or redundant network payloads in sparsely activated layers. When evaluated on an 8-H100 GPU node with MoE models comprising up to 128 experts and 16K token sequences, FlashMoE achieves up to 9x higher GPU utilization, 6x lower latency, 5.7x higher throughput, and 4x better overlap efficiency compared to state-of-the-art baselines, despite using FP32, whereas the baselines use FP16. FlashMoE shows that principled GPU kernel-hardware co-design is key to unlocking the performance ceiling of large-scale distributed ML. We provide code at https://github.com/osayamenja/FlashMoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32e34763-4ecd-4810-bd93-eb7dbeb53574Cited by top-tier papers5
- Attack of the Bubbles: Straggler-Resilient Pipeline Parallelism for Large Model TrainingTianyuan Wu, Lunxi Cao, Hanfeng Lu, Xiaoxiao Jiang et al.NSDI 2026 · 5 citations
- OpGuard: Bitwise Alignment for Precise and General Debugging of Production LLM TrainingZiming Zhou, Yinjie Zhao, Hang Zhu, Wenxiao Wang et al.OSDI 2026 · 2 citations
- MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang et al.ISCA 2026
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
- StreamEP: Straggler-Tolerant MoE Decoding without Communication BarriersYizhuo Liang, Shaoyu Wang, Jaeyong Song, Yanqi Zhou et al.SOSP 2026
Builds on16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained modelsJiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang et al.PPoPP 2022 · 97 citations
Related papers
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin et al.NeurIPS 2025 · 6 citations
- X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC PlatformsYueming Yuan, Ahan Gupta, Jianping Li, Sajal Dash et al.SC 2025 · 3 citations
- Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesXinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu et al.INFOCOM 2024 · 13 citations
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui et al.ICML 2025
