Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration
Wenju Sun, Qingyong Li, Wen Wang, Yang Liu, Yangliao Geng, Boyang Li
Abstract
Multi-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing methods primarily approach this by minimizing differences between task-specific experts and the unified model, either from a parameter-level or a task-loss perspective. However, parameter-level methods exhibit a significant performance gap compared to the upper bound, while task-loss approaches entail costly secondary training procedures. In contrast, we observe that performance degradation closely correlates with feature drift, i.e., differences in feature representations of the same sample caused by model merging. Motivated by this observation, we propose Layer-wise Optimal Task Vector Merging (LOT Merging), a technique that explicitly minimizes feature drift between task-specific experts and the unified model in a layer-by-layer manner. LOT Merging can be formulated as a convex quadratic optimization problem, enabling us to analytically derive closed-form solutions for the parameters of linear and normalization layers. Consequently, LOT Merging achieves efficient model consolidation through basic matrix operations. Extensive experiments across vision and vision-language benchmarks demonstrate that LOT Merging significantly outperforms baseline methods, achieving improvements of up to 4.4% (ViT-B/32) over state-of-the-art approaches. The source code is available at https://github.com/SunWenJu123/model-merging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Model Merging in the Essential SubspaceLonghua Li, Lei Qi, Qi Tian, Xin GengCVPR 2026 · 6 citations
- EvoGM: Learning to Merge LLMs via Evolutionary Generative OptimizationTao Jiang, Xinmeng Yu, Chenhao Yi, Yiling Wu et al.ICML 2026 · 1 citation
- SyMerge: From Non-Interference to Synergistic Merging via Single-Layer AdaptationAecheon Jung, Seunghwan Lee, Dongyoon Han, Sungeun HongICML 2026 · 1 citation
- Towards Dynamic Modality Alignment in Multimodal Continual LearningJiayao Tan, Fan Lyu, Tianle Liu, Fuyuan Hu et al.CVPR 2026
- Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core SubspaceWenju Sun, Qingyong Li, Tiancheng Li, Yangliao Geng et al.ICML 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
Related papers
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- Outlier-Aware Model Merging for Efficient Multitask InferenceQiyuan Zhu, Lujun Li, Dezhi Li, Jiacheng Liu et al.ACM MM 2025 · 2 citations
- DC-Merge: Improving Model Merging with Directional ConsistencyHan-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo, Shimin Di et al.CVPR 2026 · 12 citations
- FW-Merging: Scaling Model Merging with Frank-Wolfe OptimizationHao Mark Chen, Shell Xu Hu, Wayne Luk, Timothy M. Hospedales et al.ICCV 2025 · 6 citations
- No Task Left Behind: Isotropic Model Merging with Common and Task-Specific SubspacesDaniel Marczak, Simone Magistri, Sebastian Cygert, Bartlomiej Twardowski et al.ICML 2025
