Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering
Corentin Dancette, Rémi Cadène, Damien Teney, Matthieu Cord
Abstract
We introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statistical regularities to produce correct answers but does not actually deploy the desired behavior. There is a need to identify possible shortcuts in a dataset and assess their use before deploying a model in the real world. The research community in VQA has focused exclusively on question-based shortcuts, where a model might, for example, answer "What is the color of the sky" with "blue" by relying mostly on the question-conditional training prior and give little weight to visual evidence. We go a step further and consider multimodal shortcuts that involve both questions and images. We first identify potential shortcuts in the popular VQA v2 training set by mining trivial predictive rules such as co-occurrences of words and visual elements. We then introduce VQA-CounterExamples (VQA-CE), an evaluation protocol based on our subset of Coun-terExamples i.e. image-question-answer triplets where our rules lead to incorrect answers. We use this new evaluation in a large-scale study of existing approaches for VQA. We demonstrate that even state-of-the-art models perform poorly and that existing techniques to reduce biases are largely ineffective in this context. Our findings suggest that past work on question-based biases in VQA has only addressed one facet of a complex issue. The code for our method is available at https://github.com/ cdancette/detect-shortcuts
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54f39d4d-7a48-4992-851a-753b95b74451Cited by top-tier papers31
- Bootstrapping Multi-View Representations for Fake News DetectionQichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian et al.AAAI 2023 · 111 citations
- Look at the Variance! Efficient Black-box Explanations with Sobol-based Sensitivity AnalysisThomas Fel, Rémi Cadène, Mathieu Chalvidal, Matthieu Cord et al.NeurIPS 2021 · 100 citations
- Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional ImagesNitzan Bitton Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt et al.ICCV 2023 · 92 citations
- SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question AnsweringVipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang et al.CVPR 2022 · 41 citations
- Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningMustafa Shukor, Alexandre Ramé, Corentin Dancette, Matthieu CordICLR 2024 · 31 citations
Builds on5
- Learning from Failure: De-biasing Classifier from Biased ClassifierJun Hyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee et al.NeurIPS 2020 · 428 citations
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 136 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang et al.CVPR 2020
Related papers
- Counterfactual Vision and Language LearningEhsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi et al.CVPR 2020
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He et al.AAAI 2023 · 28 citations
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
- Which Shortcut Solution Do Question Answering Models Prefer to Learn?Kazutoshi Shinoda, Saku Sugawara, Akiko AizawaAAAI 2023 · 10 citations
- A Case Study of the Shortcut Effects in Visual Commonsense ReasoningKeren Ye, Adriana KovashkaAAAI 2021 · 47 citations
