Mitigating Parameter Interference in Model Merging via Sharpness-Aware Fine-Tuning
Yeoreum Lee, Jinwook Jung, Sungyong Baik
Abstract
Large-scale deep learning models with a pretraining-finetuning paradigm have led to a surge of numerous task-specific models fine-tuned from a common pre-trained model. Recently, several research efforts have been made on merging these large models into a single multi-task model, particularly with simple arithmetic on parameters. Such merging methodology faces a central challenge: interference between model parameters fine-tuned on different tasks. Few recent works have focused on designing a new fine-tuning scheme that can lead to small parameter interference, however at the cost of the performance of each task-specific fine-tuned model and thereby limiting that of a merged model. To improve the performance of a merged model, we note that a fine-tuning scheme should aim for (1) smaller parameter interference and (2) better performance of each fine-tuned model on the corresponding task. In this work, we aim to design a new fine-tuning objective function to work towards these two goals. In the course of this process, we find such objective function to be strikingly similar to sharpness-aware minimization (SAM) objective function, which aims to achieve generalization by finding flat minima. Drawing upon our observation, we propose to fine-tune pre-trained models via sharpness-aware minimization. The experimental and theoretical results showcase the effectiveness and orthogonality of our proposed approach, improving performance upon various merging and fine-tuning methods. Our code is available at https://github.com/baiklab/SAFT-Merge .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Train with Perturbation, Infer after Merging: A Two-Stage Framework for Continual LearningHaomiao Qiu, Miao Zhang, Ziyue Qiao, Liqiang NieNeurIPS 2025 · 8 citations
- Navigating the Accuracy-Size Trade-Off with Flexible Model MergingAkash Balasaheb Dhasade, Divyansh Jhunjhunwala, Milos Vujasinovic, Gauri Joshi et al.ICLR 2026 · 2 citations
- Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential RecommendationHuwei Ji, Jiajie Su, Yuyuan Li, Xiaohua Feng et al.KDD 2026 · 1 citation
- From Memorization to Parameter Interference: How Overtraining Experts Harms Model MergingStefan Horoi, Guy Wolf, Eugene Belilovsky, Gintare Karolina DziugaiteICML 2026
- Merge to Remember: Sharpness-Aware Isotropic Merging for Continual LearningQun Yang, Enneng Yang, Wei Chen, Li Shen et al.ICML 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
Related papers
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationHao Ban, Gokul Ram Subramani, Kaiyi JiICCV 2025 · 3 citations
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- RobustMerge: Parameter-Efficient Model Merging for MLLMs with Direction RobustnessFanhu Zeng, Haiyang Guo, Fei Zhu, Li Shen et al.NeurIPS 2025 · 28 citations
- Learn to Merge: Meta-Learning for Adaptive Multi-Task Model MergingJun Chen, Qin Zhang, Weizhi Zhang, Xiao Luo et al.ICML 2026
- Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsWenju Sun, Qingyong Li, Wen Wang, Yangliao Geng et al.ACM MM 2025 · 3 citations
