Continual Model Merging without Data: Dual Projections for Balancing Stability and Plasticity
Enneng Yang, Anke Tang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie M. Zhang
摘要
Model merging integrates multiple expert models with diverse capabilities into a unified framework, facilitating collaborative learning. However, most existing methods assume simultaneous access to all models, which is often impractical in real-world scenarios where models are received sequentially. While some studies have investigated continual model merging (CMM)-which involves sequentially merging multiple models-the challenge of balancing prior knowledge (stability) and incorporating new tasks (plasticity) remains unresolved. This paper, for the first time, formally defines the stability and plasticity of CMM from the perspective of orthogonal projection. Subsequently, we analyze the relationships among the spaces spanned by task data, historical gradients, and accumulated gradients. Building on this, we propose a data-free Dual Orthogonal Projection (DOP) method, which eliminates data dependence and mitigates interference between the merged model and models for old and new tasks by projecting their parameter differences onto their respective approximate data spaces. Finally, to solve potential conflicts between stability and plasticity, we reformulate DOP as a multi-objective optimization problem and employ a multi-gradient descent algorithm to obtain a Pareto-optimal solution. Extensive experiments across multiple architectures and task configurations validate that our approach significantly outperforms state-of-the-art CMM methods. * This work was completed during Enneng Yang's study and employment at institutions 1, 2, 5. † Corresponding authors. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
• We formalize the two key optimization objectives of CMM, namely stability and plasticity, within the framework of orthogonal projection theory, and we further examine the relationships among the subspaces spanned by task data, gradients, and accumulated gradients.
• We propose a data-free Dual Orthogonal Projection (DOP) method for CMM, recasting it as a multi-objective optimization problem to obtain a Pareto-optimal solution, ensuring an effective trade-off between stability and plasticity.
• We conduct extensive CMM experiments on four architectures (covering vision and language models) with varying task numbers, and show that our method consistently outperforms state-ofthe-art (SOTA) CMM approaches on both old and new tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- Sparsity Curse: Understanding RLVR Model Parameter Space from Model MergingChenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu 等KDD 2026 · 被引用 2 次
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 被引用 1 次
- Unlocking the Potential of Continual Model Merging: An ODE PerspectiveLihong Lin, Haidong KangICML 2026
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho 等NeurIPS 2021 · 被引用 630 次
相关 Paper
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting PlasticityZihuan Qiu, Lei Wang, Yang Cao, Runtong ZHANG 等ICLR 2026 · 被引用 4 次
- Modeling Multi-Task Model Merging as Adaptive Projective Gradient DescentYongxian Wei, Anke Tang, Li Shen, Zixuan Hu 等ICML 2025
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo 等NeurIPS 2025 · 被引用 8 次
- BECAME: Bayesian Continual Learning with Adaptive Model MergingMei Li, Yuxiang Lu, Qinyan Dai, Suizhi Huang 等ICML 2025
- ACE-Merging: Data-Free Model Merging with Adaptive Covariance EstimationBo Xu, Haotian Wu, Hehai Lin, Weiquan Huang 等CVPR 2026 · 被引用 5 次
