ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
Shaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Yu Yang, Bo Li, Xiaowen Chu
Abstract
In recent years, large-scale models can be easily scaled to trillions of parameters with sparsely activated mixture-of-experts (MoE), which significantly improves the model quality while only requiring a sub-linear increase in computational costs. However, MoE layers require the input data to be dynamically routed to a particular GPU for computing during distributed training. The highly dynamic property of data routing and high communication costs in MoE make the training system low scaling efficiency on GPU clusters. In this work, we propose an extensible and efficient MoE training system, ScheMoE, which is equipped with several features. 1) ScheMoE provides a generic scheduling framework that allows the communication and computation tasks in training MoE models to be scheduled in an optimal way. 2) ScheMoE integrates our proposed novel all-to-all collective which better utilizes intra- and inter-connect bandwidths. 3) ScheMoE supports easy extensions of customized all-to-all collectives and data compression approaches while enjoying our scheduling algorithm. Extensive experiments are conducted on a 32-GPU cluster and the results show that ScheMoE outperforms existing state-of-the-art MoE systems, Tutel and Faster-MoE, by 9%-30%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a3b89aa8-1d71-4960-8523-4638eace583bCited by top-tier papers15
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin et al.NSDI 2025 · 27 citations
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 22 citations
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXinglin Pan, Wenxiang Lin, Lin Zhang, Shaohuai Shi et al.ASPLOS 2025 · 12 citations
- PopFetcher: Towards Accelerated Mixture-of-Experts Training Via Popularity Based Expert-Wise PrefetchJunyi Zhang, Chuanhu Ma, Xiong Wang, Yuntao Nie et al.USENIX ATC 2025 · 8 citations
- HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert SwapWenxiang Lin, Xinglin Pan, Lin Zhang, Shaohuai Shi et al.INFOCOM 2026 · 7 citations
Related papers
- Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated SchedulesXinglin Pan, Wenxiang Lin, Shaohuai Shi, Xiaowen Chu et al.INFOCOM 2024 · 13 citations
- FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts TrainingYunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin et al.NeurIPS 2025 · 6 citations
- Shortcut-connected Expert Parallelism for Accelerating Mixture of ExpertsWeilin Cai, Juyong Jiang, Le Qin, Junwei Cui et al.ICML 2025
- Janus: A Unified Distributed Training Framework for Sparse Mixture-of-Experts ModelsJuncai Liu, Jessie Hui Wang, Yimin JiangSIGCOMM 2023 · 46 citations
- Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model TrainingYechan Kim, Hwijoon Lim, Dongsu HanICML 2024 · 10 citations
