DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
Yuchen Feng, Bowen Shen, Naibin Gu, Jiaxuan Zhao, Peng Fu, Zheng Lin, Weiping Wang
摘要
Large language models (LLMs) with the Mixture-of-Experts (MoE) architecture achieve high cost-efficiency by selectively activating a subset of the parameters. Despite the inference efficiency of MoE LLMs, the training of extensive experts from scratch incurs substantial overhead, whereas reconstructing a dense LLM into an MoE LLM significantly reduces the training budget. However, existing reconstruction methods often overlook the diversity among experts, leading to potential redundancy. In this paper, we come up with the observation that a specific LLM exhibits notable diversity after being pruned on different calibration datasets, based on which we present a Diversity-Enhanced reconstruction method named DIVE. The recipe of DIVE includes domain affinity mining, pruning-based expert reconstruction, and efficient retraining. Specifically, the reconstruction includes pruning and reassembly of the feed-forward network (FFN) module. After reconstruction, we efficiently retrain the model on routers, experts and normalization modules. We implement DIVE on Llama-style LLMs with open-source training corpora. Experiments show that DIVE achieves training efficiency with minimal accuracy tradeoffs, outperforming existing pruning and MoE reconstruction methods with the same number of activated parameters. Code is available at: https://github.com/yuchenblah/DIVE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal UnderstandingYuchen Feng, Zhenyu Zhang, Naibin Gu, Yilong Chen 等CVPR 2026 · 被引用 2 次
- Uncertainty-Aware Routing for Principled Alignment with MoE DynamicsYilong Chen, Junyuan Shang, Yuchen Feng, Zhenyu Zhang 等ACL 2026
- CBP-Tuning: Efficient Local Customization for Black-box Large Language ModelsJiaxuan Zhao, Naibin Gu, Yuchen Feng, Xiyu Liu 等EMNLP 2025
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
相关 Paper
- Routing Experts: Learning to Route Dynamic Experts in Existing Multi-modal Large Language ModelsQiong Wu, Zhaoxi Ke, Yiyi Zhou, Xiaoshuai Sun 等ICLR 2025
- Expert Divergence Learning for MoE-based Language ModelsJiaang Li, Haibin Chen, Langming Liu, Yujin Yuan 等ICLR 2026 · 被引用 3 次
- Breaking the Echo Chamber: A Dynamic Ensemble Pruning Perspective on MoEXinlai Kang, Dunyao Xue, Zhengbo Wang, Chengshuo Du 等ICML 2026
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language ModelsXudong Lu, Qi Liu, Yuhui Xu, Aojun Zhou 等ACL 2024 · 被引用 16 次
- Delta Decompression for MoE-based LLMs CompressionHao Gu, Wei Li, Lujun Li, Qiyuan Zhu 等ICML 2025
