Logical Implications for Visual Question Answering Consistency
Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman
Abstract
Despite considerable recent progress in Visual Question Answering (VQA) models, inconsistent or contradictory answers continue to cast doubt on their true reasoning capabilities. However, most proposed methods use indirect strategies or strong assumptions on pairs of questions and answers to enforce model consistency. Instead, we propose a novel strategy intended to improve model performance by directly reducing logical inconsistencies. To do this, we introduce a new consistency loss term that can be used by a wide range of the VQA models and which relies on knowing the logical relation between pairs of questions and answers. While such information is typically not available in VQA datasets, we propose to infer these logical relations using a dedicated language model and use these in our proposed consistency loss function. We conduct extensive experiments on the VQA Introspect and DME datasets and show that our method brings improvements to state-of-theart VQA models while being robust across different architectures and settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12105c1b-f406-4abc-affc-6dfebb73fb7eCited by top-tier papers1
Ask how each one uses itBuilds on19
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Medical Visual Question Answering via Conditional ReasoningLi-Ming Zhan, Bo Liu, Lu Fan, Jiaxin Chen et al.ACM MM 2020 · 157 citations
- TRAR: Routing the Attention Spans in Transformer for Visual Question AnsweringYiyi Zhou, Tianhe Ren, Chaoyang Zhu, Xiaoshuai Sun et al.ICCV 2021 · 128 citations
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu et al.CVPR 2022 · 115 citations
Related papers
- Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven OptimizationYue Zhang, Liqiang Jing, Vibhav GogateAAAI 2025 · 13 citations
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
- Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question AnsweringSongtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang et al.ACL 2026
- Language Models with RationalityNora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson et al.EMNLP 2023 · 7 citations
