How Transferable Are Reasoning Patterns in VQA?
Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche, Romain Vuillemot, Christian Wolf
摘要
Since its inception, Visual Question Answering (VQA) is notoriously known as a task, where models are prone to exploit biases in datasets to find shortcuts instead of performing high-level reasoning. Classical methods address this by removing biases from training data, or adding branches to models to detect and remove biases. In this paper, we argue that uncertainty in vision is a dominating factor preventing the successful learning of reasoning in vision and language problems. We train a visual oracle and in a large scale study provide experimental evidence that it is much less prone to exploiting spurious dataset biases compared to standard models. We propose to study the attention mechanisms at work in the visual oracle and compare them with a SOTA Transformer-based model. We provide an in-depth analysis and visualizations of reasoning patterns obtained with an online visualization tool which we make publicly available 1 . We exploit these insights by transferring reasoning patterns from the oracle to a SOTA Transformer-based VQA model taking standard noisy visual inputs via fine-tuning. In experiments we report higher overall accuracy, as well as accuracy on infrequent answers for each question type, which provides evidence for improved generalization and a decrease of the dependency on dataset biases. * Both authors contributed equally. 1 https://reasoningpatterns.github.io GQA data GT objects GT classes GQA data R-CNN objects R-CNN embed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringVipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang 等CVPR 2022 · 被引用 41 次
- VisQA: X-raying Vision and Language Reasoning in TransformersTheo Jaunet, Corentin Kervadec, Romain Vuillemot, Grigory Antipov 等IEEE VIS 2021 · 被引用 30 次
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski 等NeurIPS 2023 · 被引用 27 次
- COVR: A Test-Bed for Visually Grounded Compositional Generalization with Real ImagesBen Bogin, Shivanshu Gupta, Matt Gardner, Jonathan BerantEMNLP 2021 · 被引用 13 次
- Supervising the Transfer of Reasoning Patterns in VQACorentin Kervadec, Christian Wolf, Grigory Antipov, Moez Baccouche 等NeurIPS 2021 · 被引用 11 次
它引用的顶会 Paper6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl 等ICLR 2021 · 被引用 620 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
- Attention Flows: Analyzing and Comparing Attention Mechanisms in Language ModelsJoseph F. DeRose, Jiayao Wang, Matthew BergerIEEE VIS 2020 · 被引用 109 次
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi 等CVPR 2021
相关 Paper
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu 等NeurIPS 2021 · 被引用 102 次
- Roses Are Red, Violets Are Blue... but Should VQA Expect Them To?Corentin Kervadec, Grigory Antipov, Moez Baccouche, Christian WolfCVPR 2021
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- A Case Study of the Shortcut Effects in Visual Commonsense ReasoningKeren Ye, Adriana KovashkaAAAI 2021 · 被引用 47 次
- Towards Reasoning Ability in Scene Text Visual Question AnsweringQingqing Wang, Liqiang Xiao, Yue Lu, Yaohui Jin 等ACM MM 2021 · 被引用 12 次
