Vision-language Incremental Learning with Dual Class-individual Memory
Fuhai Chen, Feng Zhang, Xiaoguang Ma, Yiyi Zhou, Jiarong Liu, Xuri Ge
Abstract
The emergence of multimodal technologies has propelled Vision-Language Incremental Learning (VLIL) into a research spotlight. Current VLIL approaches predominantly inherit unimodal paradigms, failing to address fundamental distinctions between visual and linguistic modalities. Crucially, the semantic gap between images and text creates divergent learning dynamics: visual data exhibits rich, distributed information while textual representations remain explicit and compact. Consequently, textual elements align with class-specific tasks, whereas individual images inherently span multiple such tasks, creating dual bottlenecks in class-level memory allocation and scene-level knowledge transfer. To overcome these challenges, we propose DCIM (Dual Class-Individual Memory), a novel framework featuring complementary mechanisms for vision-language continual learning. For class-level constraints, our Hierarchical Class Memory Management (HCMM) strategy dynamically allocates memory resources across object categories. It employs forgetting simulation to identify and preserve the most vulnerable samples, ensuring robust long-term knowledge retention. For scene-level adaptation, the Scene Reconstruction Memory(SRM) module captures generalized environmental representations, enabling contextual transfer to novel classes and disambiguation of semantically related concepts within shared scenes.Extensive experiments on two vision-language tasks, i.e., visual question answering (VQA) and Image captioning (IC), demonstrate the effectiveness and excellent generalization ability of our approach, achieving superior performance under continual learning settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Incremental Learning Using Conditional Adversarial NetworksYe Xiang, Ying Fu, Pan Ji, Hua HuangICCV 2019 · 188 citations
- DDGR: Continual Learning with Deep Diffusion-based Generative ReplayRui Gao, Weiwei LiuICML 2023 · 101 citations
- RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningRiccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de WeijerNeurIPS 2020 · 55 citations
Related papers
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningBorui Kang, Lei Wang, Zhiping Wu, Tao Feng et al.ICCV 2025 · 4 citations
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
- Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningLinlan Huang, Xusheng Cao, Haori Lu, Yifan Meng et al.ICCV 2025 · 12 citations
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li et al.AAAI 2024 · 9 citations
