Check It Again: Progressive Visual Question Answering via Visual Entailment
Qingyi Si, Zheng Lin, Mingyu Zheng, Peng Fu, Weiping Wang
摘要
While sophisticated Visual Question Answering models have achieved remarkable success, they tend to answer questions only according to superficial correlations between question and answer. Several recent approaches have been developed to address this language priors problem. However, most of them predict the correct answer according to one best output without checking the authenticity of answers. Besides, they only explore the interaction between image and question, ignoring the semantics of candidate answers. In this paper, we propose a select-and-rerank (SAR) progressive framework based on Visual Entailment. Specifically, we first select the candidate answers relevant to the question or the image, then we rerank the candidate answers by a visual entailment task, which verifies whether the image semantically entails the synthetic statement of the question and each candidate answer. Experimental results show the effectiveness of our proposed framework, which establishes a new state-of-the-art accuracy on VQA-CP v2 with a 7.55% improvement. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun 等NeurIPS 2024 · 被引用 31 次
- Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringQingyi Si, Yuanxin Liu, Zheng Lin, Peng Fu 等EMNLP 2023 · 被引用 2 次
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin 等AAAI 2024 · 被引用 1 次
- Transform to Transfer: Boosting Adversarial Attack Transferability on Vision-Language Pre-training ModelsYang Li, Jia-Li Yin, Luojun Lin, Wei LinCVPR 2026
- Characterizing the Predictive Impact of Modalities with Supervised Latent-Variable ModelingDivyam Madaan, Sumit Chopra, Kyunghyun ChoICML 2026
它引用的顶会 Paper2
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 被引用 136 次
相关 Paper
- Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question AnsweringDongze Hao, Qunbo Wang, Longteng Guo, Jie Jiang 等EMNLP 2024 · 被引用 4 次
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu 等ICML 2026
- Re-Attention for Visual Question AnsweringWenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang 等AAAI 2020 · 被引用 90 次
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 被引用 8 次
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
