STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
Jaeseong Lee, Seung-won Hwang, Aurick Qiao, Daniel F. Campos, Zhewei Yao, Yuxiong He
Abstract
Mixture-of-experts (MoEs) have been adopted to reduce inference costs by sparsely activating experts in large language models (LLMs). Despite these reductions, the massive number of parameters in MoEs still makes them expensive to serve. Conventionally, unstructured or structured pruning has been considered to reduce the number of parameters. Our key contribution is exploring the interpolation between structured and unstructured pruning, to propose a novel structured-then-unstructured (STUN) approach outperforming both structured and unstructured pruning, especially for MoEs. In the first stage, we show a scalable expert pruning with O(1) forward pass, unlike existing work requiring O( k n √ n ) forward passes for n experts that cannot scale for recent MoEs with hundreds of experts. We then show our expert-pruned MoEs are robust to unstructured pruning to follow. Experiments on Snowflake Arctic and Mixtral show that our proposal is highly effective-For Snowflake Arctic, a 480B-sized MoE with 128 experts, our method needs only one H100 and two hours to achieve nearly no loss in performance with 40% sparsity, even in generative tasks such as GSM8K, where state-of-the-art structured or unstructured pruning methods fail. The code is publicly available. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 511b4233-b342-4566-9d05-ba239001e544Cited by top-tier papers8
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMsXiaodong Chen, Mingming Ha, Zhenzhong Lan, Jing Zhang et al.ICLR 2026 · 12 citations
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inferenceYushu Zhao, Zheng Wang, Minjia ZhangICML 2026 · 8 citations
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingLujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin et al.AAAI 2026 · 2 citations
- C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts ModelsLin Li, Yan Wang, Zhuopeng WangAAAI 2026 · 1 citation
- Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert OffloadingHanfei Yu, Xingqi Cui, Hong Zhang, Hao Wang et al.EuroSys 2026 · 1 citation
Builds on17
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 994 citations
Related papers
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language ModelsXudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou et al.ACL 2024 · 16 citations
- DOT-MoE: Differentiable Optimal Transport for MoEficationUdbhav Bamba, Arnav Chavan, Aryamaan Thakur, Steven Teig et al.ICML 2026
- Masks Can be Learned as an Alternative to ExpertsPeiyu Liu, Tianwen Wei, Bo Zhu, Xin Zhao et al.ACL 2025 · 1 citation
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoEGeng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang et al.ICLR 2026 · 14 citations
- Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and InferenceWeilin Cai, Le Qin, Shwai He, Junwei Cui et al.ICML 2026
