Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering
Feifei Zhang, Zhihao Wang, Xi Zhang, Changsheng Xu
摘要
Visual Question Answering (VQA) is a widely explored multimodal task aimed at answering questions based on images. Recently, a few studies have started to investigate continual learning in VQA to cope with evolving multimodal data streams. However, these studies fall short of tackling another critical issue in real-world VQA applications: the long-tailed distribution of data. In this paper, we introduce Continual Long-Tailed Visual Question Answering (CLT-VQA) and identify two critical challenges: innertask prototype drift, where classifier prototypes become biased toward majority classes due to imbalanced data, and inter-task feature drift, where learned features shift over time, causing forgetting of previously learned knowledge. To address these challenges, we propose a unified dualbalance approach that integrates a Balanced Classifier Prototype (BCP) learning module and a Multi-modal Feature Alignment (MFA) module. The BCP optimizes classifier prototypes to achieve balanced class representation, while the MFA aligns features consistently across tasks, preventing catastrophic forgetting. Extensive experimental results demonstrate that our method outperforms existing models, validating the effectiveness of the proposed approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper39
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan 等ICLR 2020 · 被引用 1,496 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
- The Majority Can Help the Minority: Context-rich Minority Oversampling for Long-tailed ClassificationSeulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun 等CVPR 2022 · 被引用 199 次
- Inducing Neural Collapse in Imbalanced Learning: Do We Really Need a Learnable Classifier at the End of Deep Neural Network?Yibo Yang, Shixiang Chen, Xiangtai Li, Liang Xie 等NeurIPS 2022 · 被引用 144 次
相关 Paper
- MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question AnsweringZhifei Li, Yiran Wang, Chenyi Xiong, Yujing Xia 等AAAI 2026
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question AnsweringTianyu Huai, Jie Zhou, Xingjiao Wu, Qin Chen 等CVPR 2025
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu 等CVPR 2026
- Prior-free Balanced Replay: Uncertainty-guided Reservoir Sampling for Long-Tailed Continual LearningLei Liu, Li Liu, Yawen CuiACM MM 2024 · 被引用 1 次
