Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
Zihan Wang, Deli Chen, Damai Dai, Runxin Xu, Zhuoshu Li, Yu Wu
Abstract
Parameter-efficient fine-tuning (PEFT) is crucial for customizing Large Language Models (LLMs) with constrained resources. Although there have been various PEFT methods for dense-architecture LLMs, PEFT for sparsearchitecture LLMs is still underexplored. In this work, we study the PEFT method for LLMs with the Mixture-of-Experts (MoE) architecture and the contents of this work are mainly threefold: (1) We investigate the dispersion degree of the activated experts in customized tasks, and found that the routing distribution for a specific task tends to be highly concentrated, while the distribution of activated experts varies significantly across different tasks. (2) We propose Expert-Specialized Fine-Tuning, or ESFT, which tunes the experts most relevant to downstream tasks while freezing the other experts and modules; experimental results demonstrate that our method not only improves the tuning efficiency, but also matches or even surpasses the performance of fullparameter fine-tuning. (3) We further analyze the impact of the MoE architecture on expertspecialized fine-tuning. We find that MoE models with finer-grained experts are more advantageous in selecting the combination of experts that are most relevant to downstream tasks, thereby enhancing both the training efficiency and effectiveness. Our code is available at https://github.com/deepseek-ai/ESFT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b44770a2-83d5-424b-b0ea-b0635701ecc4Cited by top-tier papers11
- SEUF: Is Unlearning One Expert Enough for Mixture-of-Experts LLMs?Haomin Zhuang, Yihua Zhang, Kehan Guo, Jinghan Jia et al.ACL 2025 · 10 citations
- AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating ProjectionsXin Yu, Yujia Wang, Jinghui Chen, Lingzhou XueNeurIPS 2025 · 8 citations
- Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE AdaptationJunzhuo Li, Bo Wang, Xiuze Zhou, Xuming HuEMNLP 2025 · 5 citations
- Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts TransformersXin Zhao, Xiaojun Chen, Bingshan Liu, Haoyu Gao et al.NeurIPS 2025 · 3 citations
- Federated Fine-Tuning of Sparsely-Activated Large Language Models on Resource-Constrained DevicesFahao Chen, Jie Wan, Peng Li, Zhou Su et al.EuroSys 2026 · 2 citations
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
Related papers
- Pushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningTed Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermis et al.ICLR 2024 · 169 citations
- LoRACoE: Improving Large Language Model via Composition-based LoRA ExpertGuanyu Li, Zhiheng Xi, Zhihao Zhang, Boyang Hong et al.EMNLP 2025
- MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language ModelsJie Cao, Tianwei Lin, Bo Yuan, Rolan Yan et al.ACL 2026 · 2 citations
- MEFT: Memory-Efficient Fine-Tuning through Sparse AdapterJitai Hao, Weiwei Sun, Xin Xin, Qi Meng et al.ACL 2024 · 4 citations
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA ExpertsYuan Zhuang, Yi Shen, Yuexin Bian, Qing Su et al.ICLR 2026 · 15 citations
