Calibrating Prompt from History for Continual Vision-Language Retrieval and Grounding
Tao Jin, Weicai Yan, Ye Wang, Sihang Cai, Qifan Shuai, Zhou Zhao
摘要
In the field of machine learning, continual learning is a crucial concept that allows models to adapt to non-stationary data distributions. However, most of the existing works focus on uni-modal settings and ignore the multi-modal data. In this paper, to enable neural networks better understand diverse modalities in real-world scenario, we investigate continual learning for two typical vision-language applications, i.e. retrieval and grounding. Instead of conventional exemplar-based methods, we leverage the pre-trained transformer model (e.g. CLIP/GLIP) and the prompt technique to tackle this problem. Under this scheme, we identify two critical limitations in existing methods: (1) Unfamiliarity across tasks, which prevents task-specific prompts from achieving forward propagation; and (2) Heterogeneity between modalities, which makes it difficult to guarantee a consistent optimization direction for prompts of different modalities. To overcome these constraints, we design Historical Prompt Calibration that includes two objectives to calibrate prompts. First, the intra-modal relevance estimation helps encode sufficient task-specific information for prompts, with the help a relevance estimator developed for recognizing task relevance. Second, the inter-modal consistency alignment enhances the agreement of the two modality-specific prompts in the current task by contrasting them with the prompts from previous tasks. We evaluate the superiority of our strategy over state-of-the arts methods by four vision-language applications, including two retrieval tasks (i.e. image- and video-text retrieval) and two grounding tasks (i.e. referring expression comprehension and segmentation).
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- Classifier-guided Gradient Modulation for Enhanced Multimodal LearningZirun Guo, Tao Jin, Jingyuan Chen, Zhou ZhaoNeurIPS 2024 · 被引用 56 次
- Chat-Driven Text Generation and Interaction for Person RetrievalZequn Xie, Chuxin Wang, Yeqiang Wang, Sihang Cai 等EMNLP 2025 · 被引用 12 次
- Low-rank Prompt Interaction for Continual Vision-Language RetrievalWeicai Yan, Ye Wang, Wang Lin, Zirun Guo 等ACM MM 2024 · 被引用 8 次
- Diff-Prompt: Diffusion-Driven Prompt Generator with Mask SupervisionWeicai Yan, Wang Lin, Zirun Guo, Ye Wang 等ICLR 2025
- DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question AnsweringBorui Li, Xingcai Zhang, Tianen Liu, Shuai Wang 等AAAI 2026
相关 Paper
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language ModelsSaurav Jha, Dong Gong, Lina YaoNeurIPS 2024 · 被引用 36 次
- LAMM: Label Alignment for Multi-Modal Prompt LearningJingsheng Gao, Jiacheng Ruan, Suncheng Xiang, Zefang Yu 等AAAI 2024 · 被引用 33 次
- Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Pengda Qin 等ICCV 2023 · 被引用 28 次
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan 等CVPR 2023
