SCoMoE: Efficient Mixtures of Experts with Structured Communication
Zhiyuan Zeng, Deyi Xiong
Abstract
Mixture-of-Experts (MoE) models are promising architectures for massively multilingual neural machine translation and large language models due to the advantage of sublinear scaling. However, the training of large MoE models is usually bottlenecked by the all-to-all communication (Lepikhin et al., 2020). To reduce the communication cost, we propose SCoMoE, an MoE architecture with structured all-to-all communication, inspired by the hierarchical architecture of the communication topology. SCoMoE encourages data to be communicated across devices through fast intra-accelerator/node communication channels, reducing communication throughput in the slow inter-node communication channel. We slice the data on the sequence dimension (SCoMoE-Seq) into three communication groups and project the data on the feature dimension (SCoMoE-Feat) into low-dimensional representations. To compensate the potential performance drop caused by the routing locality in SCoMoE, we further propose a token clustering approach to aggregating related tokens from different devices before the MoE layers. The sigmoid gating in the balanced router used in the token clustering is substituted with the softmax gating with differential sorting. Experiments on bilingual and massively multilingual machine translation demonstrate that SCoMoE achieves a speedup of 1.44x over GShard with comparable performance, and substantially outperforms Gshard (2.8 BLEU) on OPUS-100 with a speedup of 1.25x.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get dd383046-0b07-42eb-8fad-e417742f6fa7Cited by top-tier papers5
- Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesXinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu et al.INFOCOM 2024 · 13 citations
- LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive HashingXiaonan Nie, Qibin Liu, Fangcheng Fu, Shenhan Zhu et al.NeurIPS 2024 · 10 citations
- PopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchJunyi Zhang, Chuanhu Ma, Xiong Wang, Yuntao Nie et al.USENIX ATC 2025 · 8 citations
- Turn Waste into Worth: Rectifying Top-k Router of MoEZhiyuan Zeng, Qipeng Guo, Zhaoye Fei, Zhangyue Yin et al.EMNLP 2024 · 1 citation
- NetMoE: Accelerating MoE Training through Dynamic Sample PlacementXinyi Liu, Yujie Wang, Fangcheng Fu, Xupeng Miao et al.ICLR 2025
Related papers
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui et al.ICML 2025
- Gating Dropout: Communication-efficient Regularization for Sparsely Activated TransformersRui Liu, Young Jin Kim, Alexandre Muzio, Hany HassanICML 2022 · 31 citations
- ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingShaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu et al.EuroSys 2024 · 26 citations
- MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in ProductionChao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong et al.EuroSys 2026 · 5 citations
- HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert SwapWenxiang Lin, Xinglin Pan, Lin Zhang, Shaohuai Shi et al.INFOCOM 2026 · 7 citations
