Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question Answering
Zi Qian, Xin Wang, Xuguang Duan, Pengda Qin, Yuhong Li, Wenwu Zhu
摘要
In the real world, a desirable Visual Question Answering model is expected to provide correct answers to new questions and images in a continual setting (recognized as CL-VQA). However, existing works formulate CL-VQA from a vision-only or language-only perspective, and straightforwardly apply the uni-modal continual learning (CL) strategies to this multi-modal task, which is improper and suboptimal. On the one hand, such a partial formulation may result in limited evaluations. On the other hand, neglecting the interactions between modalities will lead to poor performance. To tackle these challenging issues, we propose a comprehensive formulation for CL-VQA from the perspective of multi-modal vision-language fusion. Based on our formulation, we further propose MulTi-Modal PRompt LearnIng with DecouPLing bEfore InTeraction (TRIPLET), a novel approach that builds on a pre-trained vision-language model and consists of decoupled prompts and prompt interaction strategies to capture the complex interactions between modalities. In particular, decoupled prompts contain learnable parameters that are decoupled w.r.t different aspects, and the prompt interaction strategies are in charge of modeling interactions between inputs and prompts. Additionally, we build two CL-VQA benchmarks for a more comprehensive evaluation. Extensive experiments demonstrate that our TRIPLET outperforms state-of-the-art methods in both uni-modal and multi-modal continual settings for CL-VQA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Continual Multimodal Contrastive LearningXiaohao Liu, Xiaobo Xia, See-Kiong Ng, Tat-Seng ChuaNeurIPS 2025 · 被引用 25 次
- Bisecle: Binding and Separation in Continual Learning for Video Language UnderstandingYue Tan, Xiaoqian Hu, Hao Xue, Celso de Melo 等NeurIPS 2025 · 被引用 14 次
- Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksZijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao 等NeurIPS 2024 · 被引用 11 次
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 被引用 3 次
- Empowering Large Language Model for Continual Video Question Answering with Collaborative PromptingChen Cai, Zheng Wang, Jianjun Gao, Wenyang Liu 等EMNLP 2024 · 被引用 3 次
它引用的顶会 Paper16
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- FLAVA: A Foundational Language And Vision Alignment ModelAmanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon 等CVPR 2022 · 被引用 483 次
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 被引用 397 次
相关 Paper
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- Calibrating Prompt from History for Continual Vision-Language Retrieval and GroundingTao Jin, Weicai Yan, Ye Wang, Sihang Cai 等ACM MM 2024 · 被引用 6 次
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 被引用 8 次
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan 等CVPR 2023
- Symbolic Replay: Scene Graph as Prompt for Continual Learning on VQA TaskStan Weixian Lei, Difei Gao, Jay Zhangjie Wu, Yuxuan Wang 等AAAI 2023 · 被引用 54 次
