Generalization and Scaling Laws for Mixture-of-ExpertsTransformers
Mansour ZOUBEIROU A MAYAKI
摘要
We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates active per-input capacity from routing combinatorics. By conditioning on fixed routing patterns and union-bounding across them, we derive a sup-norm covering-number bound whose metric entropy scales with the active parameter budget and incurs a MoE-specific routing overhead. Combined with a standard ERM analysis for squared loss, this yields a generalization bound under a -dimensional manifold data model and targets, showing that approximation and estimation trade off as in dense networks once active parameters are accounted for appropriately. We further prove a constructive approximation theorem for MoE architectures, showing that, under the approximation construction, error can decrease either by scaling active capacity or by increasing the number of experts, depending on the dominant bottleneck. From these results we derive neural scaling laws for model size, data size, and compute-optimal tradeoffs. Overall, our results provide a transparent statistical reference point for reasoning about MoE scaling, clarifying which behaviors are certified by worst-case theory and which must arise from data-dependent routing structure or optimization dynamics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann 等NeurIPS 2021 · 被引用 1,213 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- Scaling Laws for Fine-Grained Mixture of ExpertsJan Ludziejewski, Jakub Krajewski, Kamil Adamczewski, Maciej Pióro 等ICML 2024 · 被引用 149 次
相关 Paper
- Hyperparameter Transfer with Mixture-of-Expert LayersTianze Jiang, Blake Bordelon, Cengiz Pehlevan, Boris HaninICML 2026 · 被引用 6 次
- Joint MoE Scaling Laws: Mixture of Experts Can Be Memory EfficientJan Ludziejewski, Maciej Pióro, Jakub Krajewski, Maciej Stefaniak 等ICML 2025
- Mixture of Parrots: Experts improve memorization more than reasoningSamy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu 等ICLR 2025
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang 等ICML 2026
- On the Benefits of Learning to Route in Mixture-of-Experts ModelsNishanth Dikkala, Nikhil Ghosh, Raghu Meka, Rina Panigrahy 等EMNLP 2023 · 被引用 9 次
