Model Merging by Uncertainty-Based Gradient Matching
Nico Daheim, Thomas Möllenhoff, Edoardo M. Ponti, Iryna Gurevych, Mohammad Emtiyaz Khan
Abstract
Models trained on different datasets can be merged by a weighted-averaging of their parameters, but why does it work and when can it fail? Here, we connect the inaccuracy of weighted-averaging to mismatches in the gradients and propose a new uncertainty-based scheme to improve the performance by reducing the mismatch. The connection also reveals implicit assumptions in other schemes such as averaging, task arithmetic, and Fisher-weighted averaging. Our new method gives consistent improvements for large language models and vision transformers, both in terms of performance and robustness to hyperparameters. Code available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94b5e482-a824-4cff-812e-34f2c20ce549Cited by top-tier papers30
- Towards Modular LLMs by Building and Reusing a Library of LoRAsOleksiy Ostapenko, Zhan Su, Edoardo M. Ponti, Laurent Charlin et al.ICML 2024 · 70 citations
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- Bayesian Uncertainty for Gradient Aggregation in Multi-Task LearningIdan Achituve, Idit Diamant, Arnon Netzer, Gal Chechik et al.ICML 2024 · 14 citations
- DC-Merge: Improving Model Merging with Directional ConsistencyHan-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo, Shimin Di et al.CVPR 2026 · 12 citations
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model MergingYongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang et al.ICLR 2026 · 10 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang et al.ICLR 2026
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
- Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia et al.CVPR 2026
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesDaniel Marczak, Simone Magistri, Sebastian Cygert, Bartlomiej Twardowski et al.ICML 2025
