Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success
Luca Zhou, Bo Zhao, Rose Yu, Emanuele Rodolà
Abstract
Model merging combines knowledge from separately fine-tuned models, yet the factors driving its success remain poorly understood. While recent work treats mergeability as an intrinsic property of the models, we show with an architectureagnostic framework that it fundamentally depends on both the merging method and the partner tasks. Using L1-regularized linear optimization over a set of interpretable pairwise metrics (e.g., gradient L 2 distance), we uncover properties correlating with post-merge normalized accuracy across five merging methods. We find that the drivers of merge success vary across architectures and merging methods, overall with only a moderate agreement (64.0% average top-5 metric overlap; 79.3% sign agreement). Crucially, however, gradient alignment metrics consistently emerge as the most fundamental signals of mergeability. These findings provide a diagnostic foundation for understanding mergeability and motivate future mergeaware fine-tuning strategies. Despite its practical appeal, model merging remains highly unpredictable. Some model pairs merge seamlessly, yielding performance close to the average of their individual task
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- Model merging with SVD to tie the KnotsGeorge Stoica, Pratik Ramesh, Boglarka Ecsedi, Leshem Choshen et al.ICLR 2025
- Scalable Model Merging with Progressive Layer-wise DistillationJing Xu, Jiazheng Li, Jingzhao ZhangICML 2025
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang et al.ICLR 2026
- DC-Merge: Improving Model Merging with Directional ConsistencyHan-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo, Shimin Di et al.CVPR 2026 · 12 citations
- FairMerging: Rethinking Model Merging through the Lens of FairnessBing Liu, Xinrui Shan, Boyu Zhang, Qiankun Zhang et al.ICML 2026
