Mixture Compressor for Mixture-of-Experts LLMs Gains More
Wei Huang, Yue Liao, Jianhui Liu, Ruifei He, Haoru Tan, Shiming Zhang, Hongsheng Li, Si Liu, Xiaojuan Qi
摘要
Mixture-of-Experts large language models (MoE-LLMs) marks a significant step forward of language models, however, they encounter two critical challenges in practice: 1) expert parameters lead to considerable memory consumption and loading latency; and 2) the current activated experts are redundant, as many tokens may only require a single expert. Motivated by these issues, we investigate the MoE-LLMs and make two key observations: a) different experts exhibit varying behaviors on activation reconstruction error, routing scores, and activated frequencies, highlighting their differing importance, and b) not all tokens are equally important-only a small subset is critical. Building on these insights, we propose MC, a training-free Mixture-Compressor for MoE-LLMs, which leverages the significance of both experts and tokens to achieve an extreme compression. First, to mitigate storage and loading overheads, we introduce Pre-Loading Mixed-Precision Quantization (PMQ), which formulates the adaptive bit-width allocation as a Linear Programming (LP) problem, where the objective function balances multi-factors reflecting the importance of each expert. Additionally, we develop Online Dynamic Pruning (ODP), which identifies important tokens to retain and dynamically select activated experts for other tokens during inference to optimize efficiency while maintaining performance. Our MC integrates static quantization and dynamic pruning to collaboratively achieve extreme compression for MoE-LLMs with less accuracy loss, ensuring an optimal trade-off between performance and efficiency. Extensive experiments confirm the effectiveness of our approach. For instance, at 2.54 bits, MC compresses 76.6% of the model, with only a 3.8% average accuracy loss in eight commonsense benchmarks. During dynamic inference, we further reduce activated parameters by 15%, with a performance drop of less than 0.6%. Remarkably, MC even surpasses floating-point 13b dense LLMs with significantly smaller parameter sizes, suggesting that mixture compression in MoE-LLMs has the potential to outperform both comparable and larger dense LLMs. Our code is available at https://github.com/Aaronhuang-778/MC-MoE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compressionMike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie 等ICLR 2026 · 被引用 47 次
- QeRL: Beyond Efficiency - Quantization-enhanced Reinforcement Learning for LLMsWei Huang, Yi Ge, Shuai Yang, Yicheng Xiao 等ICLR 2026 · 被引用 19 次
- Unveiling Super Experts in Mixture-of-Experts Large Language ModelsZunhai Su, Qingyuan Li, HaoZhang, Weihao Ye 等ICLR 2026 · 被引用 16 次
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert SkippingYushi Huang, Zining Wang, Zhihang Yuan, Yifu Ding 等CVPR 2026 · 被引用 15 次
- MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoEGeng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language ModelsYuanteng Chen, Yuantian Shao, Peisong Wang, Jian ChengACL 2025
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy AnalysisYuzhuang Xu, Xu Han, Yuanchi Zhang, Yixuan Wang 等AAAI 2026 · 被引用 2 次
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language ModelsXudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou 等ACL 2024 · 被引用 16 次
- Delta Decompression for MoE-based LLMs CompressionHao Gu, Wei Li, Lujun Li, Qiyuan Zhu 等ICML 2025
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceZhixuan Chen, Xing Hu, Dawei Yang, Zukang Xu 等ICML 2025
