DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
Sikai Bai, Haoxi Li, Jie Zhang, Zicong Hong, Song Guo
Abstract
Despite the significant breakthrough of Mixture-of-Experts (MoE), the increasing scale of these MoE models presents huge memory and storage challenges. Existing MoE pruning methods, which involve reducing parameter size with a uniform sparsity across all layers, often lead to suboptimal outcomes and performance degradation due to varying expert redundancy in different MoE layers. To address this, we propose a non-uniform pruning strategy, dubbed Differentiable Expert Pruning (DiEP), which adaptively adjusts pruning rates at the layer level while jointly learning inter-layer importance, effectively capturing the varying redundancy across different MoE layers. By transforming the global discrete search space into a continuous one, our method handles exponentially growing non-uniform expert combinations, enabling adaptive gradient-based pruning. Extensive experiments on five advanced MoE models demonstrate the efficacy of our method across various NLP tasks. Notably, DiEP retains around 92% of original performance on Mixtral 87B with only half the experts, outperforming other pruning methods by up to 7.1% on the challenging MMLU dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Stabilizing Differentiable Architecture Search via Perturbation-based RegularizationXiangning Chen, Cho-Jui HsiehICML 2020 · 235 citations
Related papers
- Less Token, More Signal: MoE Expert Pruning via Critical Token SelectionZeliang Zong, Kai Zhang, Yarong Wang, wenming tan et al.ICML 2026
- C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts ModelsLin Li, Yan Wang, Zhuopeng WangAAAI 2026 · 1 citation
- Mosaic Pruning: A Hierarchical Framework for Generalizable Pruning of Mixture-of-Experts ModelsWentao Hu, Mingkuan Zhao, Shuangyong Song, Xiaoyan Zhu et al.AAAI 2026 · 4 citations
- HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output SpaceKe Li, Zheng Yang, Zhongbin Zhou, Xuefeng et al.ICLR 2026 · 4 citations
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoEGeng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang et al.ICLR 2026 · 14 citations
