Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural Networks
Mohammed Nowaz Rabbani Chowdhury, Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen
Abstract
In deep learning, mixture-of-experts (MoE) activates one or few experts (sub-networks) on a per-sample or per-token basis, resulting in significant computation reduction. The recently proposed patch-level routing in MoE (pMoE) divides each input into patches (or tokens) and sends patches () to each expert through prioritized routing. pMoE has demonstrated great empirical success in reducing training and inference costs while maintaining test accuracy. However, the theoretical explanation of pMoE and the general MoE remains elusive. Focusing on a supervised classification task using a mixture of two-layer convolutional neural networks (CNNs), we show for the first time that pMoE provably reduces the required number of training samples to achieve desirable generalization (referred to as the sample complexity) by a factor in the polynomial order of , and outperforms its single-expert counterpart of the same or even larger capacity. The advantage results from the discriminative routing property, which is justified in both theory and practice that pMoE routers can filter label-irrelevant patches and route similar class-discriminative patches to the same expert. Our experimental results on MNIST, CIFAR-10, and CelebA support our theoretical findings on pMoE's generalization and show that pMoE can avoid learning spurious correlations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- On the Convergence and Sample Complexity Analysis of Deep Q-Networks with ε-Greedy ExplorationShuai Zhang, Hongkang Li, Meng Wang, Miao Liu et al.NeurIPS 2023 · 57 citations
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-ExpertsHeming Zou, Yunliang Zang, Wutong Xu, Yao Zhu et al.NeurIPS 2025 · 38 citations
- How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui et al.ICML 2024 · 37 citations
- DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-IdentificationYuhao Wang, Yang Liu, Aihua Zheng, Pingping ZhangAAAI 2025 · 31 citations
- What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional EncodingHongkang Li, Meng Wang, Tengfei Ma, Sijia Liu et al.ICML 2024 · 23 citations
Builds on7
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- When Do Neural Networks Outperform Kernel Methods?Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, Andrea MontanariNeurIPS 2020 · 217 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 104 citations
Related papers
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-ExpertsMohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang et al.ICML 2024 · 18 citations
- Robust Mixture-of-Expert Training for Convolutional Neural NetworksYihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang et al.ICCV 2023 · 43 citations
- Rethinking Convergence in MoE Training: The Role of Routing SparsityWeihao Zhu, Long Shi, Kang Wei, Zhe Wang et al.ICML 2026
- Teacher-Guided Routing for Sparse Vision Mixture-of-ExpertsMasahiro Kada, Ryota Yoshihashi, Satoshi Ikehata, Rei Kawakami et al.CVPR 2026
