Lune

ICML2026Top-tier venue

Rethinking Convergence in MoE Training: The Role of Routing Sparsity

Weihao Zhu, Long Shi, Kang Wei, Zhe Wang, Yipeng Zhou, Haixia Zhang

2026Year

Abstract

In Mixture-of-Experts (MoE) training, sparse routing, i.e., activating only the top-KK experts per token, is essential for balancing convergence speed and computational cost. However, existing works typically choose KK empirically, without theoretical guidance. To address this gap, we characterize the convergence behavior of MoE training using stochastic optimization theory. Specifically, we derive a convergence upper bound of O(1+M/KT)\mathcal{O}\left(\frac{1+M/K}{\sqrt{T}}\right), where TT is the number of training iterations and MM is the total number of experts per MoE layer. This result guarantees convergence and shows that increasing KK can accelerate training. By further fixing the total computational budget RR (in FLOPs), we obtain a refined bound of O(KR+MKR)\mathcal{O}\left(\sqrt{\frac{K}{R}} + \frac{M}{\sqrt{K R}}\right), which is convex in KK and implies the existence of an optimal K∗∈[1,M]K^{*}\in[1,M] that achieves the best convergence performance. Extensive experiments validate our theoretical analysis under diverse settings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 2eb7bf78-92b8-4115-9280-6123b1f3cc4c

Builds on16

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines