Check It Again: Progressive Visual Question Answering via Visual Entailment
Qingyi Si, Zheng Lin, Mingyu Zheng, Peng Fu, Weiping Wang
Abstract
While sophisticated Visual Question Answering models have achieved remarkable success, they tend to answer questions only according to superficial correlations between question and answer. Several recent approaches have been developed to address this language priors problem. However, most of them predict the correct answer according to one best output without checking the authenticity of answers. Besides, they only explore the interaction between image and question, ignoring the semantics of candidate answers. In this paper, we propose a select-and-rerank (SAR) progressive framework based on Visual Entailment. Specifically, we first select the candidate answers relevant to the question or the image, then we rerank the candidate answers by a visual entailment task, which verifies whether the image semantically entails the synthetic statement of the question and each candidate answer. Experimental results show the effectiveness of our proposed framework, which establishes a new state-of-the-art accuracy on VQA-CP v2 with a 7.55% improvement. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6ffc70c6-8eab-4fe9-a7d1-91e3b2fc1f29Cited by top-tier papers6
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun et al.NeurIPS 2024 · 31 citations
- Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringQingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu et al.EMNLP 2023 · 2 citations
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2024 · 1 citation
- Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training ModelsYang Li, Jia-Li Yin, Luojun Lin, Wei LinCVPR 2026
- Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable ModelingDivyam Madaan, Sumit Chopra, Kyunghyun ChoICML 2026
Builds on2
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 136 citations
Related papers
- Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question AnsweringDongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang et al.EMNLP 2024 · 4 citations
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu et al.ICML 2026
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang et al.AAAI 2020 · 90 citations
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 8 citations
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
