SCHUR-A*: Layer-wise Optimal Expert Pruning for MoEs via Schur-Complement Guided A* Search
Zheng Chen, Weifeng Yang, Jianxiao Tang, Buhui Yao
摘要
Sparse Mixture-of-Experts (MoE) language models enable conditional computation but face deployment challenges due to the memory wall: while few experts are activated per token, the entire model must reside in memory. Existing expert pruning methods primarily rely on independent ranking, failing to account for the complex inter-dependencies and redundancies between experts. In this paper, we formulate post-training MoE pruning as a reconstruction-driven subset selection problem, aiming to minimize layer-output distortion under a cardinality constraint. We introduce SCHUR-A*, an algorithm that leverages A* search to achieve globally optimal expert selection within each layer. To maintain computational tractability, we derive a novel, admissible heuristic upper bound using a Schur-complement-based relaxation of the reconstruction objective. This tight bound allows for aggressive pruning of the search space while mathematically guaranteeing optimality. Furthermore, we propose an automated strategy to balance fidelity and memory reduction across heterogeneous layers via knee-point detection. Extensive experiments on Qwen3-30B-A3B demonstrate that SCHUR-A* significantly outperforms greedy and ranking-based baselines, maintaining comparable performance even under aggressive pruning ratios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Mixture-of-Experts with Expert Choice RoutingYanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du 等NeurIPS 2022 · 被引用 933 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI ScaleSamyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang 等ICML 2022 · 被引用 523 次
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compressionMike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie 等ICLR 2026 · 被引用 47 次
相关 Paper
- Attribution-Guided and Coverage-Maximized Pruning for Structural MoE CompressionYifu Ding, jiacheng wang, Ge Yang, Yongcheng Jing 等ICML 2026
- Discovering Important Experts for Mixture-of-Experts Models Pruning Through a Theoretical PerspectiveWeizhong Huang, Yuxin Zhang, Xiawu Zheng, Fei Chao 等NeurIPS 2025 · 被引用 12 次
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoEGeng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang 等ICLR 2026 · 被引用 14 次
- Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information DensityZhendong Mi, Yixiao Chen, Pu Zhao, Xiaodong Yu 等ICML 2026 · 被引用 6 次
- HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output SpaceKe Li, Zheng Yang, Zhongbin Zhou, Xuefeng 等ICLR 2026 · 被引用 4 次
