Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language Tasks
Zijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao, Huaimin Wang
Abstract
Continual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent efforts focus on devising specific frameworks and various replay strategies, striving for a typical learning-forgetting trade-off. Surprisingly, both our empirical research and theoretical analysis demonstrate that the stability of the model in consecutive zero-shot predictions serves as a reliable indicator of its anti-forgetting capabilities for previously learned tasks. Motivated by these insights, we develop a novel replay-free CL method named ZAF (Zero-shot Antidote to Forgetting), which preserves acquired knowledge through a zero-shot stability regularization applied to wild data in a plug-and-play manner. To enhance efficiency in adapting to new tasks and seamlessly access historical models, we introduce a parameter-efficient EMA-LoRA neural architecture based on the Exponential Moving Average (EMA). ZAF utilizes new data for low-rank adaptation (LoRA), complemented by a zero-shot antidote on wild data, effectively decoupling learning from forgetting. Our extensive experiments demonstrate ZAF's superior performance and robustness in pre-trained models across various continual VL concept learning tasks, achieving leads of up to 3.70%, 4.82%, and 4.38%, along with at least a 10x acceleration in training speed on three benchmarks, respectively. Additionally, our zero-shot antidote significantly reduces forgetting in existing models by at least 6.37%. Our code is available at https://github.com/Zi-Jian-Gao/Stabilizing-Zero-Shot-Prediction-ZAF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45a09626-cf7b-4fc2-8c0a-8e275812481fCited by top-tier papers10
- Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsDingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen et al.EMNLP 2025 · 4 citations
- Pansharpening for Thin-Cloud Contaminated Remote Sensing Images: A Unified Framework and Benchmark DatasetSongcheng Du, Yang Zou, Jiaxin Li, Mingxuan Liu et al.AAAI 2026 · 3 citations
- Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal LearningJiayu Zhang, Chuangxin Zhao, Canran Xiao, Ruibo Duan et al.ICLR 2026
- CoMem: Compositional Concept-Graph Memory for Vision-Language AdaptationHeng Zhou, Jing Tang, Jusheng Zhang, Yanshu Li et al.ICLR 2026
- Decouple Your Discovery and Memory in Continual Generalized Category DiscoveryJiawei Yu, Zijian Gao, Xingxing Zhang, Xuan Liu et al.CVPR 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Decomposing and Composing: Towards Efficient Vision-Language Continual Learning via Rank-1 Expert Pool in a Single LoRAZhan Fa, Yue Duan, Jian Zhang, Lei Qi et al.AAAI 2026 · 1 citation
- ConStruct-VL: Data-Free Continual Structured VL Concepts LearningJames Seale Smith, Paola Cascante-Bonilla, Assaf Arbelle, Donghyun Kim et al.CVPR 2023
- Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language ModelsYujin Jo, Taesup KimICLR 2026 · 2 citations
- Revisiting Weight Regularization for Low-Rank Continual LearningYaoyue Zheng, Yin Zhang, Joost van de Weijer, Gido van de Ven et al.ICLR 2026 · 7 citations
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
