Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, Yongbin Li
Abstract
In this paper, we unveil that Language Models (LMs) can acquire new capabilities by assimilating parameters from homologous models without retraining or GPUs. We first introduce DARE to set most delta parameters (i.e., the disparity between fine-tuned and pre-trained parameters) to zeros without affecting the abilities of Supervised Fine-Tuning (SFT) LMs, which randomly Drops delta parameters with a ratio And REscales the remaining ones by to approximate the original embeddings. Then, we use DARE as a versatile plug-in to sparsify delta parameters of multiple SFT homologous models for mitigating parameter interference and merge them into a single model by parameter fusing. We experiment with encoder- and decoder-based LMs, showing that: (1) SFT delta parameter value ranges are typically small (within 0.002) with extreme redundancy, and DARE can effortlessly eliminate 90% or even 99% of them; (2) DARE can merge multiple task-specific LMs into one LM with diverse capabilities. Notably, this phenomenon is more pronounced in large-scale LMs, where the merged LM reveals the potential to surpass the performance of any source LM, providing a new discovery. We also utilize DARE to create a merged LM that ranks first among models with 7 billion parameters on the Open LLM Leaderboard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers267
- MiniCache: KV Cache Compression in Depth Dimension for Large Language ModelsAkide Liu, Jing Liu, Zizheng Pan, Yefei He et al.NeurIPS 2024 · 160 citations
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- EMR-Merging: Tuning-Free High-Performance Model MergingChenyu Huang, Peng Ye, Tao Chen, Tong He et al.NeurIPS 2024 · 134 citations
- Representation Surgery for Multi-Task Model MergingEnneng Yang, Li Shen, Zhenyi Wang, Guibing Guo et al.ICML 2024 · 96 citations
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
Builds on11
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Layer-adaptive Sparsity for the Magnitude-based PruningJaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn et al.ICLR 2021 · 331 citations
Related papers
- Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li et al.ACL 2026
- Training-free LLM Merging for Multi-task LearningZichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang et al.ACL 2025
- Knowledge Fusion of Large Language ModelsFanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan et al.ICLR 2024 · 113 citations
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMsZixuan Ren, Jinliang Lu, Junhong Wu, Yang Zhao et al.ICLR 2026 · 2 citations
