Continual Model Merging without Data: Dual Projections for Balancing Stability and Plasticity
Enneng Yang, Anke Tang, Li Shen, Guibing Guo, Xingwei Wang, Xiaochun Cao, Jie M. Zhang
Abstract
Model merging integrates multiple expert models with diverse capabilities into a unified framework, facilitating collaborative learning. However, most existing methods assume simultaneous access to all models, which is often impractical in real-world scenarios where models are received sequentially. While some studies have investigated continual model merging (CMM)-which involves sequentially merging multiple models-the challenge of balancing prior knowledge (stability) and incorporating new tasks (plasticity) remains unresolved. This paper, for the first time, formally defines the stability and plasticity of CMM from the perspective of orthogonal projection. Subsequently, we analyze the relationships among the spaces spanned by task data, historical gradients, and accumulated gradients. Building on this, we propose a data-free Dual Orthogonal Projection (DOP) method, which eliminates data dependence and mitigates interference between the merged model and models for old and new tasks by projecting their parameter differences onto their respective approximate data spaces. Finally, to solve potential conflicts between stability and plasticity, we reformulate DOP as a multi-objective optimization problem and employ a multi-gradient descent algorithm to obtain a Pareto-optimal solution. Extensive experiments across multiple architectures and task configurations validate that our approach significantly outperforms state-of-the-art CMM methods. * This work was completed during Enneng Yang's study and employment at institutions 1, 2, 5. † Corresponding authors. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
• We formalize the two key optimization objectives of CMM, namely stability and plasticity, within the framework of orthogonal projection theory, and we further examine the relationships among the subspaces spanned by task data, gradients, and accumulated gradients.
• We propose a data-free Dual Orthogonal Projection (DOP) method for CMM, recasting it as a multi-objective optimization problem to obtain a Pareto-optimal solution, ensuring an effective trade-off between stability and plasticity.
• We conduct extensive CMM experiments on four architectures (covering vision and language models) with varying task numbers, and show that our method consistently outperforms state-ofthe-art (SOTA) CMM approaches on both old and new tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 3 citations
- Sparsity Curse: Understanding RLVR Model Parameter Space from Model MergingChenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu et al.KDD 2026 · 2 citations
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 1 citation
- Unlocking the Potential of Continual Model Merging: An ODE PerspectiveLihong Lin, Haidong KangICML 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- SWAD: Domain Generalization by Seeking Flat MinimaJunbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho et al.NeurIPS 2021 · 630 citations
Related papers
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Stability, Promoting PlasticityZihuan Qiu, Lei Wang, Yang Cao, Runtong ZHANG et al.ICLR 2026 · 4 citations
- Modeling Multi-Task Model Merging as Adaptive Projective Gradient DescentYongxian Wei, Anke Tang, Li Shen, Zixuan Hu et al.ICML 2025
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo et al.NeurIPS 2025 · 8 citations
- BECAME: Bayesian Continual Learning with Adaptive Model MergingMei Li, Yuxiang Lu, Qinyan Dai, Suizhi Huang et al.ICML 2025
- ACE-Merging: Data-Free Model Merging with Adaptive Covariance EstimationBo Xu, Haotian Wu, Hehai Lin, Weiquan Huang et al.CVPR 2026 · 5 citations
