Continual Vision-Language Representation Learning with Off-Diagonal Information
Zixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang, Qi Tian
摘要
Large-scale multi-modal contrastive learning frameworks like CLIP typically require a large amount of image-text samples for training. However, these samples are always collected continuously in real scenarios. This paper discusses the feasibility of continual CLIP training using streaming data. Unlike continual learning based on self-supervised learning methods for pure images, which is empirically robust against catastrophic forgetting, CLIP's performance degeneration in the continual setting is significant and non-neglectable. By analyzing the changes in the model's representation space during continual CLIP training from a spatial geometry perspective, we explore and summarize these spatial variations as Spatial Disorder (SD), which can be divided into Intra-modal Rotation and Inter-modal Deviation. Moreover, we empirically and theoretically demonstrate how SD leads to a performance decline for CLIP on cross-modal retrieval tasks. To alleviate SD, we propose a new continual vision-language representation learning framework Mod-X: Maintain off-diagonal information-matriX. By selectively aligning the off-diagonal information distribution of contrastive matrices, the Mod-X improves the capability of the multi-modal model by maintaining the multi-modal representation space alignment on the old data domain during continuously fitting the new training data domain. Experiments on commonly used datasets with different scales and scopes have demonstrated the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- A Unified Approach to Domain Incremental Learning with Memory: Theory and AlgorithmHaizhou Shi, Hao WangNeurIPS 2023 · 被引用 60 次
- GraphControl: Adding Conditional Control to Universal Graph Pre-trained Models for Graph Domain Transfer LearningYun Zhu, Yaoke Wang, Haizhou Shi, Zhenshuo Zhang 等WWW 2024 · 被引用 43 次
- Degeneration-Tuning: Using Scrambled Grid shield Unwanted Concepts from Stable DiffusionZixuan Ni, Longhui Wei, Jiacheng Li, Siliang Tang 等ACM MM 2023 · 被引用 13 次
- Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksZijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao 等NeurIPS 2024 · 被引用 11 次
- Continual Vision-Language Retrieval via Dynamic Knowledge RectificationZhenyu Cui, Yuxin Peng, Xun Wang, Manyu Zhu 等AAAI 2024 · 被引用 11 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
相关 Paper
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
- Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningLinlan Huang, Xusheng Cao, Haori Lu, Yifan Meng 等ICCV 2025 · 被引用 12 次
- Subspace Alignment for CLIP-based Continual Learning via Canonical Correlation AnalysisHuan Zhang, Shuyu Dong, Yujin Zheng, Dingwen Wang 等CVPR 2026
