Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent
Yongxian Wei, Anke Tang, Li Shen, Zixuan Hu, Chun Yuan, Xiaochun Cao
Abstract
Merging multiple expert models offers a promising approach for performing multi-task learning without accessing their original data. Existing methods attempt to alleviate task conflicts by sparsifying task vectors or promoting orthogonality among them. However, they overlook the fundamental target of model merging: the merged model performs as closely as possible to taskspecific models on respective tasks. We find these methods inevitably discard task-specific information that, while causing conflicts, is crucial for performance. Based on our findings, we frame model merging as a constrained optimization problem (i.e., minimizing the gap between the merged model and individual models, subject to the constraint of retaining shared knowledge) and solve it via adaptive projective gradient descent. Specifically, we align the merged model with individual models by decomposing and reconstituting the loss function, alleviating conflicts through data-free optimization of task vectors. To retain shared knowledge, we optimize this objective by projecting gradients within a shared subspace spanning all tasks. Moreover, we view merging coefficients as adaptive learning rates and propose a task-aware, training-free strategy. Experiments show that our plug-andplay approach consistently outperforms previous methods, achieving state-of-the-art results across diverse architectures and tasks in both vision and NLP domains. Our code is available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 285d55a4-15cd-405c-b634-2f6afc88c7c3Cited by top-tier papers33
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- AdaRank: Adaptive Rank Pruning for Enhanced Model MergingChanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim et al.ICLR 2026 · 14 citations
- Continual Model Merging without Data: Dual Projections for Balancing Stability and PlasticityEnneng Yang, Anke Tang, Li Shen, Guibing Guo et al.NeurIPS 2025 · 13 citations
- DC-Merge: Improving Model Merging with Directional ConsistencyHan-Chen Zhang, Zi-Hao Zhou, Mao-Lin Luo, Shimin Di et al.CVPR 2026 · 12 citations
- Merge before Forget: A Single LoRA Continual Learning via Continual MergingFuli Qiao, Mehrdad MahdaviICLR 2026 · 11 citations
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
Related papers
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- ACE-Merging: Data-Free Model Merging with Adaptive Covariance EstimationBo Xu, Haotian Wu, Hehai Lin, Weiquan Huang et al.CVPR 2026 · 5 citations
- Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core SubspaceWenju Sun, Qingyong Li, Tiancheng Li, Yangliao Geng et al.ICML 2026
- Whoever Started the interference Should End It: Guiding Data-Free Model Merging via Task VectorsRunxi Cheng, Feng Xiong, Yongxian Wei, Wanyun Zhu et al.ICML 2025
- SyMerge: From Non-Interference to Synergistic Merging via Single-Layer AdaptationAecheon Jung, Seunghwan Lee, Dongyoon Han, Sungeun HongICML 2026 · 1 citation
