Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu
摘要
Large language models (LLMs) have demonstrated considerable proficiency in general natural language processing (NLP) tasks. Instruction tuning, a successful paradigm, enhances the ability of LLMs to follow natural language instructions and exhibit robust generalization across general tasks. However, these models often encounter performance limitations across multiple tasks due to constrained model capacity. Expanding this capacity during the instruction tuning phase poses significant challenges. To address this issue, we introduce parameter-efficient sparsity crafting (PESC), which crafts dense models into sparse models using the mixture-of-experts (MoE) architecture. PESC integrates adapters into the MoE layers of sparse models, differentiating experts without altering the individual weights within these layers. This method significantly reduces computational costs and GPU memory requirements, facilitating model capacity expansion through a minimal parameter increase when guaranteeing the quality of approximation in function space compared to original sparse upcycling. Our empirical evaluation demonstrates the effectiveness of the PESC method. Using PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.5. Our code is available at https://github.com/wuhy68/ Parameter-Efficient-MoE .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu 等ACL 2026 · 被引用 6 次
- Sparse Spectral LoRA: Routed Experts for Medical VLMsOmid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao 等CVPR 2026 · 被引用 3 次
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-ExpertsYifeng Ding, Jiawei Liu, Yuxiang Wei, Lingming ZhangACL 2024 · 被引用 2 次
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationYouwei Zheng, Yuxi Ren, Xin Xia, Xuefeng Xiao 等ICCV 2025 · 被引用 1 次
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 被引用 1 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter MergingTingfeng Hui, Zhenyu Zhang, Shuohuan Wang, Yu Sun 等ACL 2025
- Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language ModelsSheng Shen, Le Hou, Yanqi Zhou, Nan Du 等ICLR 2024 · 被引用 87 次
- Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-ExpertsShengzhuang Chen, Ying Wei, Jonathan Richard SchwarzACL 2025
- MEFT: Memory-Efficient Fine-Tuning through Sparse AdapterJitai Hao, Weiwei Sun, Xin Xin, Qi Meng 等ACL 2024 · 被引用 4 次
- Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsZihan Wang, Deli Chen, Damai Dai, Runxin Xu 等EMNLP 2024 · 被引用 2 次
