MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
Chao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong, Juncai Liu, Xiang Li, Ningxin Zheng, Xi Wang, Cong Xie, Qi Huang, Wen Heng, Yiyuan Ma
Abstract
We present MegaScale-MoE, a production system tailored for the efficient training of large-scale mixture-of-experts (MoE) models. MoE emerges as a promising architecture to scale large language models (LLMs) to unprecedented sizes, thereby enhancing model performance. However, existing MoE training systems experience a degradation in training efficiency, exacerbated by the escalating scale of MoE models and the continuous evolution of hardware.
Recognizing the pivotal role of efficient communication in enhancing MoE training, MegaScale-MoE customizes communication-efficient parallelism strategies for attention and FFNs in each MoE layer and adopts a holistic approach to overlap communication with computation at both inter-and intra-operator levels. Additionally, MegaScale-MoE applies communication compression with adjusted communication patterns to lower precision, further improving training efficiency. When training a 352B MoE model on 1,440 NVIDIA Hopper GPUs, MegaScale-MoE achieves a training throughput of 1.41M tokens/s, improving the efficiency by 1.88× compared to Megatron-LM. We share our operational experience in accelerating MoE training and hope that by offering our insights in system design, this work will motivate future research in MoE systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a653db84-0f65-48b1-9996-937c4b83fc76Cited by top-tier papers10
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 3 citations
- SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA OffloadingXingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang et al.NSDI 2026 · 2 citations
- UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsYipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng et al.SIGCOMM 2026
- LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts TrainingXinyi Liu, Yujie Wang, Fangcheng Fu, Xuefeng Xiao et al.ASPLOS 2026
- UEP: Portable Expert-Parallel CommunicationZiming Mao, Yihan Zhang, Chihan Cui, Zhen Huang et al.OSDI 2026
Builds on20
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase et al.USENIX ATC 2021 · 657 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
Related papers
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert ParallelismRuidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu et al.SIGCOMM 2025 · 19 citations
- BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and InferenceZewen Jin, Shengnan Wang, Jiaan Zhu, Hongrui Zhan et al.AAAI 2025 · 6 citations
- HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert SwapWenxiang Lin, Xinglin Pan, Lin Zhang, Shaohuai Shi et al.INFOCOM 2026 · 7 citations
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
- FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE PipeliningGuichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu et al.ACL 2025
