Overcoming Language Priors in VQA via Decomposed Linguistic Representations
Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, Qi Wu
Abstract
Most existing Visual Question Answering (VQA) models overly rely on language priors between questions and answers. In this paper, we present a novel method of language attention-based VQA that learns decomposed linguistic representations of questions and utilizes the representations to infer answers for overcoming language priors. We introduce a modular language attention mechanism to parse a question into three phrase representations: type representation, object representation, and concept representation. We use the type representation to identify the question type and the possible answer set (yes/no or specific concepts such as colors or numbers), and the object representation to focus on the relevant region of an image. The concept representation is verified with the attended region to infer the final answer. The proposed method decouples the language-based concept discovery and vision-based concept verification in the process of answer inference to prevent language priors from dominating the answering process. Experiments on the VQA-CP dataset demonstrate the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8404fe97-d548-43c3-9a10-8c855bca3b79Cited by top-tier papers24
- DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language ModelsGe Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou et al.NeurIPS 2023 · 252 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
- Greedy Gradient Ensemble for Robust Visual Question AnsweringXinzhe Han, Shuhui Wang, Chi Su, Qingming Huang et al.ICCV 2021 · 94 citations
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 74 citations
Builds on1
Related papers
- Separating Skills and Concepts for Novel Visual Question AnsweringSpencer Whitehead, Hui Wu, Heng Ji, Rogério Feris et al.CVPR 2021
- TOA: Task-oriented Active VQAXiaoying Xing, Mingfu Liang, Ying WuNeurIPS 2023 · 20 citations
- Focal and Composed Vision-semantic Modeling for Visual Question AnsweringYudong Han, Yangyang Guo, Jianhua Yin, Meng Liu et al.ACM MM 2021 · 14 citations
- Aligned Dual Channel Graph Convolutional Network for Visual Question AnsweringQingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng et al.ACL 2020 · 79 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
