Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
Lujun Li, Qiyuan Zhu, Jiacheng Wang, Xiaoyu Qin, Wei Li, Hao Gu, Sirui Han, Yike Guo
Abstract
Mixture of Experts (MoE) LLMs face significant obstacles due to their massive parameter scale, which imposes memory, storage, and deployment challenges. Although recent expert merging methods aim to achieve greater efficiency by consolidating several experts, they are fundamentally hindered by parameter conflicts arising from expert specialization. In this paper, we present Sub-MoE, a novel MoE compression framework via Subspace Expert Merging. Our key insight is to perform joint Singular Value Decomposition (SVD) on concatenated expert weights, reducing conflicting parameters by extracting shared U -matrices while enabling effective merging of the expert-specific V components. Specifically, Sub-MoE consists of two innovative stages: (1) Adaptive Expert Clustering, which groups functionally coherent experts via K-means clustering based on cosine similarity of expert outputs; and (2) Subspace Expert Merging, which first performs Experts Union Decomposition to derive the shared U -matrix across experts in the same group, then applies frequency-based merging for individual V -matrices, and completes expert reconstruction using the merged V -matrix. In this way, we align and fuse experts in a shared subspace. Additionally, the framework can be extended with intraexpert compression for further inference optimization. Extensive experiments on Mixtral, DeepSeek, and Qwen-1.5/3 MoE LLMs demonstrate that our Sub-MoE significantly outperforms existing expert pruning and merging methods. Notably, our Sub-MoE maintains 96%/86% of original performance with 25%/50% expert reduction on Mixtral-8×7B in zero-shot benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0693037e-d7d5-432b-a9b9-42429bdcd75bCited by top-tier papers10
- PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inferenceYushu Zhao, Zheng Wang, Minjia ZhangICML 2026 · 8 citations
- KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language ModelsZukang Xu, Zhixiong Zhao, Xing Hu, Zhixuan Chen et al.ICLR 2026 · 7 citations
- Effective MoE-based LLM Compression by Exploiting Heterogeneous Inter-Group Experts Routing Frequency and Information DensityZhendong Mi, Yixiao Chen, Pu Zhao, Xiaodong Yu et al.ICML 2026 · 6 citations
- HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output SpaceKe Li, Zheng Yang, Zhongbin Zhou, Xuefeng et al.ICLR 2026 · 4 citations
- CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy AnalysisYuzhuang Xu, Xu Han, Yuanchi Zhang, Yixuan Wang et al.AAAI 2026 · 2 citations
Builds on22
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- MoE-SVD: Structured Mixture-of-Experts LLMs Compression via Singular Value DecompositionWei Li, Lujun Li, Hao Gu, You-Liang Huang et al.ICML 2025
- Retraining-free Merging of Sparse MoE via Hierarchical ClusteringI-Chun Chen, Hsu-Shen Liu, Wei-Fang Sun, Chen-Hao Chao et al.ICML 2025
- Delta Decompression for MoE-based LLMs CompressionHao Gu, Wei Li, Lujun Li, Qiyuan Zhu et al.ICML 2025
- TD-MoE: Tensor Decomposition for MoE ModelsYuebin XU, YANHONG WANG, Xuemei Peng, Hui Zang et al.ICLR 2026
- EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language ModelsYuanteng Chen, Yuantian Shao, Peisong Wang, Jian ChengACL 2025
