Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering
Corentin Dancette, Rémi Cadène, Damien Teney, Matthieu Cord
摘要
We introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statistical regularities to produce correct answers but does not actually deploy the desired behavior. There is a need to identify possible shortcuts in a dataset and assess their use before deploying a model in the real world. The research community in VQA has focused exclusively on question-based shortcuts, where a model might, for example, answer "What is the color of the sky" with "blue" by relying mostly on the question-conditional training prior and give little weight to visual evidence. We go a step further and consider multimodal shortcuts that involve both questions and images. We first identify potential shortcuts in the popular VQA v2 training set by mining trivial predictive rules such as co-occurrences of words and visual elements. We then introduce VQA-CounterExamples (VQA-CE), an evaluation protocol based on our subset of Coun-terExamples i.e. image-question-answer triplets where our rules lead to incorrect answers. We use this new evaluation in a large-scale study of existing approaches for VQA. We demonstrate that even state-of-the-art models perform poorly and that existing techniques to reduce biases are largely ineffective in this context. Our findings suggest that past work on question-based biases in VQA has only addressed one facet of a complex issue. The code for our method is available at https://github.com/ cdancette/detect-shortcuts
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Bootstrapping Multi-View Representations for Fake News DetectionQichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian 等AAAI 2023 · 被引用 111 次
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity AnalysisThomas Fel, Rémi Cadène, Mathieu Chalvidal, Matthieu Cord 等NeurIPS 2021 · 被引用 100 次
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt 等ICCV 2023 · 被引用 92 次
- SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringVipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang 等CVPR 2022 · 被引用 41 次
- Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningMustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu CordICLR 2024 · 被引用 31 次
它引用的顶会 Paper5
- Learning from Failure: De-biasing Classifier from Biased ClassifierJun Hyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee 等NeurIPS 2020 · 被引用 428 次
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha 等NeurIPS 2020 · 被引用 163 次
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 被引用 136 次
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 被引用 111 次
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang 等CVPR 2020
相关 Paper
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi 等CVPR 2020
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He 等AAAI 2023 · 被引用 28 次
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu 等NeurIPS 2021 · 被引用 102 次
- Which Shortcut Solution Do Question Answering Models Prefer to Learn?Kazutoshi Shinoda, Saku Sugawara, Akiko AizawaAAAI 2023 · 被引用 10 次
- A Case Study of the Shortcut Effects in Visual Commonsense ReasoningKeren Ye, Adriana KovashkaAAAI 2021 · 被引用 47 次
