A Case Study of the Shortcut Effects in Visual Commonsense Reasoning
Keren Ye, Adriana Kovashka
Abstract
Visual reasoning and question-answering have gathered attention in recent years. Many datasets and evaluation protocols have been proposed; some have been shown to contain bias that allows models to ``cheat'' without performing true, generalizable reasoning. A well-known bias is dependence on language priors (frequency of answers) resulting in the model not looking at the image. We discover a new type of bias in the Visual Commonsense Reasoning (VCR) dataset. In particular we show that most state-of-the-art models exploit co-occurring text between input (question) and output (answer options), and rely on only a few pieces of information in the candidate options, to make a decision. Unfortunately, relying on such superficial evidence causes models to be very fragile. To measure fragility, we propose two ways to modify the validation data, in which a few words in the answer choices are modified without significant changes in meaning. We find such insignificant changes cause models' performance to degrade significantly. To resolve the issue, we propose a curriculum-based masking approach, as a mechanism to perform more robust training. Our method improves the baseline by requiring it to pay attention to the answers as a whole, and is more effective than prior masking strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6918c476-26ff-4dc5-b486-694711eae7a3Cited by top-tier papers10
- Bootstrapping Multi-View Representations for Fake News DetectionQichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian et al.AAAI 2023 · 111 citations
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He et al.AAAI 2023 · 28 citations
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen et al.ACM MM 2023 · 11 citations
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding et al.ACL 2023 · 10 citations
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando et al.ICML 2024 · 8 citations
Builds on6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun et al.AAAI 2021 · 414 citations
Related papers
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen et al.ACM MM 2021 · 22 citations
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng et al.CVPR 2020
- Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?Ming Liu, Hao Chen, Jindong Wang, Liwen Wang et al.AAAI 2026
- How Transferable Are Reasoning Patterns in VQA?Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche et al.CVPR 2021
