PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Yushu Zhao, Zheng Wang, Minjia Zhang
Abstract
Mixture-of-Experts (MoE) have shown strong potential in scaling language models efficiently by activating only a small subset of experts per input. However, their deployment remains limited due to the high memory overhead associated with storing all expert parameters, particularly as the number of experts increases. To address this challenge, prior works have explored expert dropping and merging strategies; however, they often suffer from notable performance drop especially at high compression ratios due to their reliance on coarsegrained tensor-or expert-level operations. In this paper, we introduce PuzzleMoE, the first MoE merging method to enable fine-grained elementwise merging while achieving both high accuracy and inference speed, via two key innovations: First, PuzzleMoE performs sparse expert merging by identifying element-wise weight redundancy and specialization. It introduces a dual-mask approach to capture both shared and expert-specific salient parameters. Second, to avoid the overhead of storing masks and signs, we introduce a bitpacked encoding scheme that reuses underutilized exponent bits, enabling efficient MoE inference on GPUs. Extensive experiments demonstrate that PuzzleMoE outperforms prior MoE compression methods by up to 16.7% on MMLU at 50% compression ratio, and achieves up to 1.80× endto-end inference throughput gain. The code is available at https://github.com/Supercomputing- System-AI-Lab/PuzzleMoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 444ec319-aecf-4ff1-aa32-59a490a0e571Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- SqueezeLLM: Dense-and-Sparse QuantizationSehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong et al.ICML 2024 · 306 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language ModelsXudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou et al.ACL 2024 · 16 citations
- NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMsHaeun Lee, Omin Kwon, Yeonhong Park, Jae W. LeeNeurIPS 2025 · 5 citations
Related papers
- D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingHaodong Wang, Qihua Zhou, Zicong Hong, Song GuoMobiCom 2025 · 8 citations
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language ModelsYuanteng Chen, Yuantian Shao, Peisong Wang, Jian ChengACL 2025
- Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert MergingLujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin et al.AAAI 2026 · 2 citations
- MoE-SVD: Structured Mixture-of-Experts LLMs Compression via Singular Value DecompositionWei Li, Lujun Li, Hao Gu, You-Liang Huang et al.ICML 2025
- CasMoE: A Cascaded Framework for Efficient MoE Inference on Resource-constrained DevicesChengcheng Wang, Haowen He, Liang Zhao, Xiaoheng Deng et al.AAAI 2026
