Generalization and Scaling Laws for Mixture-of-ExpertsTransformers
Mansour ZOUBEIROU A MAYAKI
Abstract
We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates active per-input capacity from routing combinatorics. By conditioning on fixed routing patterns and union-bounding across them, we derive a sup-norm covering-number bound whose metric entropy scales with the active parameter budget and incurs a MoE-specific routing overhead. Combined with a standard ERM analysis for squared loss, this yields a generalization bound under a -dimensional manifold data model and targets, showing that approximation and estimation trade off as in dense networks once active parameters are accounted for appropriately. We further prove a constructive approximation theorem for MoE architectures, showing that, under the approximation construction, error can decrease either by scaling active capacity or by increasing the number of experts, depending on the dominant bottleneck. From these results we derive neural scaling laws for model size, data size, and compute-optimal tradeoffs. Overall, our results provide a transparent statistical reference point for reasoning about MoE scaling, clarifying which behaviors are certified by worst-case theory and which must arise from data-dependent routing structure or optimization dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c28d2de8-a9a7-4dc4-9731-3d06f14af39bBuilds on9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang et al.ICML 2022 · 523 citations
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal et al.ICML 2021 · 382 citations
- Scaling Laws for Fine-Grained Mixture of ExpertsJan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro et al.ICML 2024 · 149 citations
Related papers
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 6 citations
- Joint MoE Scaling Laws: Mixture of Experts Can Be Memory EfficientJan Ludziejewski, Maciej Pióro, Jakub Krajewski, Maciej Stefaniak et al.ICML 2025
- Mixture of Parrots: Experts improve memorization more than reasoningSamy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu et al.ICLR 2025
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang et al.ICML 2026
- On the Benefits of Learning to Route in Mixture-of-Experts ModelsNishanth Dikkala, Nikhil Ghosh, Raghu Meka, Rina Panigrahy et al.EMNLP 2023 · 9 citations
