From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question Answering
Mingrui Lao, Yanming Guo, Yu Liu, Wei Chen, Nan Pu, Michael S. Lew
Abstract
Most Visual Question Answering (VQA) models are faced with language bias when learning to answer a given question, thereby failing to understand multimodal knowledge simultaneously. Based on the fact that VQA samples with different levels of language bias contribute differently for answer prediction, in this paper, we overcome the language prior problem by proposing a novel Language Bias driven Curriculum Learning (LBCL) approach, which employs an easy-to-hard learning strategy with a novel difficulty metric Visual Sensitive Coefficient (VSC). Specifically, in the initial training stage, the VQA model mainly learns the superficial textual correlations between questions and answers (easy concept) from more-biased examples, and then progressively focuses on learning the multimodal reasoning (hard concept) from less-biased examples in the following stages. The curriculum selection of examples on different stages is according to our proposed difficulty metric VSC, which is to evaluate the difficulty driven by the language bias of each VQA sample. Furthermore, to avoid the catastrophic forgetting of the learned concept during the multi-stage learning procedure, we propose to integrate knowledge distillation into the curriculum learning framework. Extensive experiments show that our LBCL can be generally applied to common VQA baseline models, and achieves remarkably better performance on the VQA-CP v1 and v2 datasets, with an overall 20% accuracy boost over baseline models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7c6802e8-d03f-4f74-ba31-82dfc2795d74Cited by top-tier papers9
- Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural NetworksNan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. GerasICML 2022 · 124 citations
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He et al.AAAI 2023 · 28 citations
- Curriculum Multi-Negative Augmentation for Debiased Video GroundingXiaohan Lan, Yitian Yuan, Hong Chen, Xin Wang et al.AAAI 2023 · 26 citations
- Understanding Unimodal Bias in Multimodal Deep Linear NetworksYedi Zhang, Peter E. Latham, Andrew M. SaxeICML 2024 · 20 citations
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 8 citations
Related papers
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2024 · 1 citation
- Intra- and Inter-Modal Curriculum for Multimodal LearningYuwei Zhou, Xin Wang, Hong Chen, Xuguang Duan et al.ACM MM 2023 · 28 citations
- Multi-Domain Lifelong Visual Question Answering via Self-Critical DistillationMingrui Lao, Nan Pu, Yu Liu, Zhun Zhong et al.ACM MM 2023 · 5 citations
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- Overcoming Language Priors in VQA via Decomposed Linguistic RepresentationsChenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia et al.AAAI 2020 · 115 citations
