EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
Yuanteng Chen, Yuantian Shao, Peisong Wang, Jian Cheng
Abstract
Mixture-of-Experts (MoE) has demonstrated promising potential in scaling LLMs. However, it is hindered by two critical challenges: (1) substantial GPU memory consumption to load all experts; (2) low activated parameters cannot be equivalently translated into inference acceleration effects. In this work, we propose EAC-MoE, an Expert-Selection Aware Compressor for MoE-LLMs, which deeply aligns with the characteristics of MoE from the perspectives of quantization and pruning, and introduces two modules to address these two challenges respectively: (1) The expert selection bias caused by low-bit quantization is a major factor contributing to the performance degradation in MoE-LLMs. Based on this, we propose Quantization with Expert-Selection Calibration (QESC), which mitigates the expert selection bias by calibrating the routers within the MoE; (2) There are always certain experts that are not crucial for the corresponding tasks, yet causing inference latency. Therefore, we propose Pruning based on Expert-Selection Frequency (PESF), which significantly improves inference speed by pruning less frequently used experts for current task. Extensive experiments demonstrate that our approach significantly reduces memory usage and improves inference speed with minimal performance degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e3804979-3be6-43c0-9450-d5a77022cd55Cited by top-tier papers7
- DartQuant: Efficient Rotational Distribution Calibration for LLM QuantizationYuantian Shao, Yuanteng Chen, Peisong Wang, Jianlin Yu et al.NeurIPS 2025 · 20 citations
- KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language ModelsZukang Xu, Zhixiong Zhao, Xing Hu, Zhixuan Chen et al.ICLR 2026 · 7 citations
- GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMsJianing Deng, Song Wang, Dongwei Wang, Zijie Liu et al.ICML 2026 · 3 citations
- TileQ: Efficient Low-Rank Quantization of Mixture-of-Experts with 2D TilingHongyaoxing Gu, Xinzhe Chen, LIJUAN HU, Liu fangfangICML 2026
- UNITE: Universal kNowledge Integration from Task-specific ExpertsShuxia Lin, Qiufeng Wang 00002, Xu Yang, Xin GengICLR 2026
Builds on7
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
- Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-VerificationAojun Zhou, Ke Wang, Zimu Lu, Weikang Shi et al.ICLR 2024 · 206 citations
- BiLLM: Pushing the Limit of Post-Training Quantization for LLMsWei Huang, Yangdong Liu, Haotong Qin, Ying Li et al.ICML 2024 · 161 citations
- Towards Accurate Post-training Network Quantization via Bit-Split and StitchingPeisong Wang, Qiang Chen, Xiangyu He, Jian ChengICML 2020 · 159 citations
Related papers
- Mixture Compressor for Mixture-of-Experts LLMs Gains MoreWei Huang, Yue Liao, Jianhui Liu, Ruifei He et al.ICLR 2025
- CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesChengcheng Wang, Haowen He, Liang Zhao, Xiaoheng Deng et al.AAAI 2026
- MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware ExpertsWei Tao, Haocheng Lu, Xiaoyang Qu, Bin Zhang et al.ACL 2025 · 8 citations
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceZhixuan Chen, Xing Hu, Dawei Yang, Zukang Xu et al.ICML 2025
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inferenceYushu Zhao, Zheng Wang, Minjia ZhangICML 2026 · 8 citations
