Continual Vision-Language Retrieval via Dynamic Knowledge Rectification
Zhenyu Cui, Yuxin Peng, Xun Wang, Manyu Zhu, Jiahuan Zhou
摘要
The recent large-scale pre-trained models like CLIP have aroused great concern in vision-language tasks. However, when required to match image-text data collected in a streaming manner, namely Continual Vision-Language Retrieval (CVRL), their performances are still limited due to the catastrophic forgetting of the learned old knowledge. To handle this issue, advanced methods are proposed to distill the affinity knowledge between images and texts from the old model to the new one for anti-forgetting. Unfortunately, existing approaches neglect the impact of incorrect affinity, which prevents the balance between the anti-forgetting of old knowledge and the acquisition of new knowledge. Therefore, we propose a novel framework called Dynamic Knowledge Rectification (DKR) that simultaneously achieves incorrect knowledge filtering and rectification. Specifically, we first filter the incorrect affinity knowledge calculated by the old model on the new data. Then, a knowledge rectification method is designed to rectify the incorrect affinities while preserving the correct ones. In particular, for the new data that can only be correctly retrieved by the new model, we rectify them with the corresponding new affinity to protect them from negative transfer. Additionally, for those that can not be retrieved by either the old or the new model, we introduce paired ground-truth labels to promote the acquisition of both old and new knowledge. Extensive experiments on several benchmark datasets demonstrate the effectiveness of our DKR and its superiority against state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksZijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao 等NeurIPS 2024 · 被引用 11 次
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen 等CVPR 2026 · 被引用 7 次
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
- Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal LearningJiayu Zhang, Chuangxin Zhao, Canran Xiao, Ruibo Duan 等ICLR 2026
- CoMem: Compositional Concept-Graph Memory for Vision-Language AdaptationHeng Zhou, Jing Tang, Jusheng Zhang, Yanshu Li 等ICLR 2026
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- Few-Shot Class-Incremental Learning via Relation Knowledge DistillationSonglin Dong, Xiaopeng Hong, Xiaoyu Tao, Xinyuan Chang 等AAAI 2021 · 被引用 215 次
相关 Paper
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
- AttriCLIP: A Non-Incremental Learner for Incremental Knowledge LearningRunqi Wang, Xiaoyue Duan, Guoliang Kang, Jianzhuang Liu 等CVPR 2023
- Continual Vision-Language Representation Learning with Off-Diagonal InformationZixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang 等ICML 2023 · 被引用 40 次
- CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language ModelsSaurav Jha, Dong Gong, Lina YaoNeurIPS 2024 · 被引用 36 次
