Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian Wolf
摘要
Models for Visual Question Answering (VQA) are notorious for their tendency to rely on dataset biases, as the large and unbalanced diversity of questions and concepts involved and tends to prevent models from learning to "reason", leading them to perform "educated guesses" instead. In this paper, we claim that the standard evaluation metric, which consists in measuring the overall in-domain accuracy, is misleading. Since questions and concepts are unbalanced, this tends to favor models which exploit subtle training set statistics. Alternatively, naively introducing artificial distribution shifts between train and test splits is also not completely satisfying. First, the shifts do not reflect real-world tendencies, resulting in unsuitable models; second, since the shifts are handcrafted, trained models are specifically designed for this particular setting, and do not generalize to other configurations. We propose the GQA-OOD benchmark designed to overcome these concerns: we measure and compare accuracy over both rare and frequent question-answer pairs, and argue that the former is better suited to the evaluation of reasoning abilities, which we experimentally validate with models trained to more or less exploit biases. In a large-scale study involving 7 VQA models and 3 bias reduction techniques, we also experimentally demonstrate that these models fail to address questions involving infrequent concepts and provide recommendations for future directions of research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Adversarial VQA: A New Benchmark for Evaluating the Robustness of VQA ModelsLinjie Li, Jie Lei, Zhe Gan, Jingjing LiuICCV 2021 · 被引用 99 次
- SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringVipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang 等CVPR 2022 · 被引用 41 次
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question AnsweringJie Ma, Min Hu, Pinghui Wang, Wangchun Sun 等NeurIPS 2024 · 被引用 31 次
- VisQA: X-raying Vision and Language Reasoning in TransformersTheo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov 等IEEE VIS 2021 · 被引用 30 次
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski 等NeurIPS 2023 · 被引用 27 次
它引用的顶会 Paper3
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin 等ICCV 2019 · 被引用 288 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
- How Transferable Are Reasoning Patterns in VQA?Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche 等CVPR 2021
相关 Paper
- When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and ApproachFeifei Zhang, Zhaoyi Zhang, Xi Zhang, Changsheng XuAAAI 2025
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng 等CVPR 2020
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual LearningZijian Gao, Zicheng Sun, Xingxing Zhang, Kele Xu 等CVPR 2026
- V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive MatricesDamien Teney, Peng Wang, Jiewei Cao, Lingqiao Liu 等AAAI 2020 · 被引用 37 次
