RCMoE: A Communication-Efficient Random Compression Framework for Resource-Constrained Mixture-of-Experts Training
Donglei Wu, Xiao Cai, Jinglei Tan, Jinda Jia, Guangming Tan, Dingwen Tao, Wen Xia, Zhihong Tian
摘要
Mixture-of-Experts (MoE) architecture with experts parallelism scales LLMs efficiently by activating only a subset of experts per input, avoiding proportional training costs. However, the intensive and heterogeneous communication substantially hinders the efficiency and scalability of MoE training in the resource-constrained scenario. Existing communication compression techniques fall short in MoE training due to: (i) Intensive training amplifies compression overhead, compromising training efficiency; (ii) Accumulated compression errors propagate through the network, degrading training quality. In this paper, we propose RCMoE, a communication-efficient Random Compression framework for MoE training with two core modules: (1) Local-Stochastic Quantization compresses the all-to-all communication by stochastically quantizing each row of the expert's intermediate computing results in parallel, effectively improving the compression efficiency and reducing compression error; (2) Probabilistic Thresholding Sparsification compresses the all-reduce communication by probabilistically sampling large gradients at high probability, thereby reducing the computational complexity and maintaining the convergence efficiency. Experiments on four typical MoE training tasks prove that RCMoE achieves higher 5.9x-8.1x total communication compression ratios and 1.3x-10.1x training speedup compared with the state-of-the-art compression techniques while maintaining the MoE training accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Stochastic Controlled Averaging for Federated Learning with Communication CompressionXinmeng Huang, Ping Li, Xiaoyun LiICLR 2024 · 被引用 288 次
- Accelerating Distributed MoE Training and Inference with LinaJiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang 等USENIX ATC 2023 · 被引用 191 次
相关 Paper
- BirdMoE: Reducing Communication Costs for Mixture-of-Experts Training Using Load-Aware Bi-random QuantizationDonglei Wu, Weihao Yang, Xiangyu Zou, Jinda Jia 等DAC 2025 · 被引用 1 次
- LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive HashingXiaonan Nie, Qibin Liu, Fangcheng Fu, Shenhan Zhu 等NeurIPS 2024 · 被引用 10 次
- ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingShaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu 等EuroSys 2024 · 被引用 26 次
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong 等EuroSys 2026 · 被引用 5 次
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin 等NeurIPS 2025 · 被引用 6 次
