Rethinking Convergence in MoE Training: The Role of Routing Sparsity
Weihao Zhu, Long Shi, Kang Wei, Zhe Wang, Yipeng Zhou, Haixia Zhang
Abstract
In Mixture-of-Experts (MoE) training, sparse routing, i.e., activating only the top- experts per token, is essential for balancing convergence speed and computational cost. However, existing works typically choose empirically, without theoretical guidance. To address this gap, we characterize the convergence behavior of MoE training using stochastic optimization theory. Specifically, we derive a convergence upper bound of , where is the number of training iterations and is the total number of experts per MoE layer. This result guarantees convergence and shows that increasing can accelerate training. By further fixing the total computational budget (in FLOPs), we obtain a refined bound of , which is convex in and implies the existence of an optimal that achieves the best convergence performance. Extensive experiments validate our theoretical analysis under diverse settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2eb7bf78-92b8-4115-9280-6123b1f3cc4cBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- On the Convergence of FedAvg on Non-IID DataXiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang et al.ICLR 2020 · 2,930 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
Related papers
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 27 citations
- Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural NetworksMohammed Nowaz Rabbani Chowdhury, Shuai Zhang, Meng Wang, Sijia Liu et al.ICML 2023 · 45 citations
- Tight Clusters Make Specialized ExpertsStefan K. Nielsen, Rachel S. Y. Teo, Laziz U. Abdullaev, Tan Minh NguyenICLR 2025
- ProbMoE: Differentiable Probabilistic Routing for Mixture-of-ExpertsHeng Zhao, Zilei Shao, Guy Van den Broeck, Zhe ZengICML 2026
