How does the optimizer implicitly bias the model merging loss landscape?
Chenxiang Zhang, Alexander Theus, Damien Teney, Antonio Orvieto, Jun Pang, Sjouke Mauw
Abstract
Model merging combines independent solutions with different capabilities into a single one while maintaining the same inference cost. Two popular approaches are linear interpolation, which simply averages multiple model weights, and task arithmetic, which combines task vectors obtained by the difference between finetuned and base models. While useful in practice, what properties make merging effective are poorly understood. This paper explores how the optimization dynamics affect the loss landscape geometry and its impact on merging success. We show that a single quantity -- the effective noise scale -- unifies the impact of different optimizer components on model merging. Across architectures and datasets, merging success is a non-monotonic function of the effective noise scale, with a distinct optimum. Decomposing this quantity, we find that larger learning rates, stronger weight decay, smaller batch sizes, and data augmentation all independently modulate the effective noise scale and exhibit the same qualitative trend. Unlike prior work connecting optimizer noise to the flatness or generalization of individual minima, we show that it also affects the global loss landscape, predicting when independently trained solutions can be successfully merged. Our findings broaden the understanding of how optimization shapes the loss landscape geometry and its consequences for model merging, suggesting that training dynamics could be further manipulated to improve model merging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb6fa672-6e84-43a8-93ed-b7bd5981abe8Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
Related papers
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang et al.ICLR 2026
- Demystifying Mergeability: Interpretable Properties to Predict Model Merging SuccessLuca Zhou, Bo Zhao, Rose Yu, Emanuele RodolàICML 2026 · 3 citations
- Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-TuningYeoreum Lee, Jinwook Jung, Sungyong BaikICLR 2025
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
- Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-TrainingWenJie Zhou, Bohan Wang, Hongtao Zhang, Chenxi Jia et al.ICML 2026
