A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
Mohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang, Pin-Yu Chen, Christopher D. Carothers
Abstract
The sparsely gated mixture of experts (MoE) architecture sends different inputs to different subnetworks (experts), through trainable routers. MoE reduces the training computation significantly for large models, but its deployment can be still memory/computation expensive for some downstream tasks. Model pruning is a popular approach to reduce inference computation, but its application in MoE architecture is largely unexplored. To the best of our knowledge, this paper provides the first provably efficient technique for pruning experts in fine-tuned MoE models. We theoretically prove that prioritizing the pruning of the experts with a smaller change of the router's l 2 norm from the pre-trained model guarantees the preservation of test accuracy, while significantly reducing the model size and the computational requirements. Although our theoretical analysis is centered on binary classification tasks on simplified MoE architecture, our expert pruning method is verified on large vision MoE models such as V-MoE and E 3 -MoE fine-tuned on benchmark datasets such as CIFAR-10, CIFAR-100, and Ima-geNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6894281f-e0dd-433f-bc1c-be47c36a5d35Cited by top-tier papers7
- How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?Hongkang Li, Meng Wang, Songtao Lu, Xiaodong Cui et al.ICML 2024 · 37 citations
- Unveiling Super Experts in Mixture-of-Experts Large Language ModelsZunhai Su, Qingyuan Li, HaoZhang, Weihao Ye et al.ICLR 2026 · 16 citations
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao et al.NeurIPS 2025 · 12 citations
- Efficient Quantization of Mixture-of-Experts with Theoretical Generalization GuaranteesMohammed Nowaz Rabbani Chowdhury, Kaoutar El Maghraoui, Hsinyu Tsai, Naigang Wang et al.ICLR 2026 · 2 citations
- QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime ReconfigurationHamid Reza Imani, Jiaxin Peng, Peiman Mohseni, Abdolah Amirany et al.ICML 2025
Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du et al.NeurIPS 2022 · 933 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
Related papers
- Teacher-Guided Routing for Sparse Vision Mixture-of-ExpertsMasahiro Kada, Ryota Yoshihashi, Satoshi Ikehata, Rei Kawakami et al.CVPR 2026
- Scaling Vision with Sparse Mixture of ExpertsCarlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann et al.NeurIPS 2021 · 1,213 citations
- Robust Mixture-of-Expert Training for Convolutional Neural NetworksYihua Zhang, Ruisi Cai, Tianlong Chen, Guanhua Zhang et al.ICCV 2023 · 43 citations
- Patch-level Routing in Mixture-of-Experts is Provably Sample-efficient for Convolutional Neural NetworksMohammed Nowaz Rabbani Chowdhury, Shuai Zhang, Meng Wang, Sijia Liu et al.ICML 2023 · 45 citations
- MoEC: Mixture of Expert ClustersYuan Xie, Shaohan Huang, Tianyu Chen, Furu WeiAAAI 2023 · 27 citations
