A Case Study of the Shortcut Effects in Visual Commonsense Reasoning
Keren Ye, Adriana Kovashka
摘要
Visual reasoning and question-answering have gathered attention in recent years. Many datasets and evaluation protocols have been proposed; some have been shown to contain bias that allows models to ``cheat'' without performing true, generalizable reasoning. A well-known bias is dependence on language priors (frequency of answers) resulting in the model not looking at the image. We discover a new type of bias in the Visual Commonsense Reasoning (VCR) dataset. In particular we show that most state-of-the-art models exploit co-occurring text between input (question) and output (answer options), and rely on only a few pieces of information in the candidate options, to make a decision. Unfortunately, relying on such superficial evidence causes models to be very fragile. To measure fragility, we propose two ways to modify the validation data, in which a few words in the answer choices are modified without significant changes in meaning. We find such insignificant changes cause models' performance to degrade significantly. To resolve the issue, we propose a curriculum-based masking approach, as a mechanism to perform more robust training. Our method improves the baseline by requiring it to pay attention to the answers as a whole, and is more effective than prior masking strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Bootstrapping Multi-View Representations for Fake News DetectionQichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian 等AAAI 2023 · 被引用 111 次
- COCA: COllaborative CAusal Regularization for Audio-Visual Question AnsweringMingrui Lao, Nan Pu, Yu Liu, Kai He 等AAAI 2023 · 被引用 28 次
- Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCRZhenyang Li, Yangyang Guo, Kejie Wang, Xiaolin Chen 等ACM MM 2023 · 被引用 11 次
- A Multi-Modal Context Reasoning Approach for Conditional Inference on Joint Textual and Visual CluesYunxin Li, Baotian Hu, Xinyu Chen, Yuxin Ding 等ACL 2023 · 被引用 10 次
- Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionIshaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando 等ICML 2024 · 被引用 8 次
它引用的顶会 Paper6
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong 等AAAI 2020 · 被引用 966 次
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu 等NeurIPS 2020 · 被引用 561 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- ERNIE-ViL: Knowledge Enhanced Vision-Language Representations through Scene GraphsFei Yu, Jiji Tang, Weichong Yin, Yu Sun 等AAAI 2021 · 被引用 414 次
相关 Paper
- From Superficial to Deep: Language Bias driven Curriculum Learning for Visual Question AnsweringMingrui Lao, Yanming Guo, Yu Liu, Wei Chen 等ACM MM 2021 · 被引用 22 次
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng 等CVPR 2020
- Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?Ming Liu, Hao Chen, Jindong Wang, Liwen Wang 等AAAI 2026
- How Transferable Are Reasoning Patterns in VQA?Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche 等CVPR 2021
