Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core Subspace
Wenju Sun, Qingyong Li, Tiancheng Li, Yangliao Geng, Albert Boyang Li
Abstract
Model merging offers an efficient solution for integrating task-specific knowledge from multiple fine-tuned models. Most existing approaches focus on manipulating the difference vectors between fine-tuned and pretrained weights, often overlooking the generalization capabilities inherent in the pretrained parameters. In this work, we revisit the role of pretrained weights in model merging and investigate their efficacy from a subspace perspective. We find that the components of pretrained weights residing in the core subspace-defined by the dominant singular vectors-are essential for maintaining generalization across diverse tasks. Specifically, we present empirical evidence that pretrained weights are nearly first-order stationary and exhibit predominantly non-negative curvature within this core subspace with respect to multi-task loss landscapes, indicating near-optimality. These findings suggest that task-specific adaptations should be injected primarily into the orthogonal complement of the core subspace, thereby preserving the generalization properties of the pretrained model. Extensive experiments on vision and vision-language tasks show that this subspace-aware strategy consistently yields improvements over state-of-the-art training-free merging methods, including Task Arithmetic, LOT Merging, ISO, and TSV. The source code is available at https://github. com/SunWenJu123/model-merging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a332da55-5643-4c08-89d3-46163b630918Builds on27
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- Localizing Task Information for Improved Model Merging and CompressionKe Wang, Nikolaos Dimitriadis, Guillermo Ortiz-Jiménez, François Fleuret et al.ICML 2024 · 107 citations
Related papers
- No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesDaniel Marczak, Simone Magistri, Sebastian Cygert, Bartlomiej Twardowski et al.ICML 2025
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- When Shared Knowledge Hurts: Spectral Over-Accumulation in Model MergingYayuan Li, Ze Peng, Jian Zhang, Jintao Guo et al.ICML 2026 · 5 citations
- Model Merging in the Essential SubspaceLonghua Li, Lei Qi, Qi Tian, Xin GengCVPR 2026 · 6 citations
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationWenju Sun, Qingyong Li, Wen Wang, Yang Liu et al.NeurIPS 2025 · 20 citations
