Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks
Haoyuan Wu, Haisheng Zheng, Zhuolun He, Bei Yu
Abstract
Large language models (LLMs) have demonstrated considerable proficiency in general natural language processing (NLP) tasks. Instruction tuning, a successful paradigm, enhances the ability of LLMs to follow natural language instructions and exhibit robust generalization across general tasks. However, these models often encounter performance limitations across multiple tasks due to constrained model capacity. Expanding this capacity during the instruction tuning phase poses significant challenges. To address this issue, we introduce parameter-efficient sparsity crafting (PESC), which crafts dense models into sparse models using the mixture-of-experts (MoE) architecture. PESC integrates adapters into the MoE layers of sparse models, differentiating experts without altering the individual weights within these layers. This method significantly reduces computational costs and GPU memory requirements, facilitating model capacity expansion through a minimal parameter increase when guaranteeing the quality of approximation in function space compared to original sparse upcycling. Our empirical evaluation demonstrates the effectiveness of the PESC method. Using PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.5. Our code is available at https://github.com/wuhy68/ Parameter-Efficient-MoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ededebbb-0c27-4377-9d54-b337d2985f6dCited by top-tier papers8
- Analytical FFN-to-MoE Restructuring via Activation Pattern AnalysisZehua Pei, Hui-Ling Zhen, Lancheng Zou, Xianzhi Yu et al.ACL 2026 · 6 citations
- Sparse Spectral LoRA: Routed Experts for Medical VLMsOmid Nejati Manzari, Hojat Asgariandehkordi, Taha Koleilat, Yiming Xiao et al.CVPR 2026 · 3 citations
- XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-ExpertsYifeng Ding, Jiawei Liu, Yuxiang Wei, Lingming ZhangACL 2024 · 2 citations
- Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image GenerationYouwei Zheng, Yuxi Ren, Xin Xia, Xuefeng Xiao et al.ICCV 2025 · 1 citation
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 1 citation
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
Related papers
- Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter MergingTingfeng Hui, Zhenyu Zhang, Shuohuan Wang, Yu Sun et al.ACL 2025
- Mixture-of-Experts Meets Instruction Tuning: A Winning Combination for Large Language ModelsSheng Shen, Le Hou, Yanqi Zhou, Nan Du et al.ICLR 2024 · 87 citations
- Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-ExpertsShengzhuang Chen, Ying Wei, Jonathan Richard SchwarzACL 2025
- MEFT: Memory-Efficient Fine-Tuning through Sparse AdapterJitai Hao, Weiwei Sun, Xin Xin, Qi Meng et al.ACL 2024 · 4 citations
- Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language ModelsZihan Wang, Deli Chen, Damai Dai, Runxin Xu et al.EMNLP 2024 · 2 citations
