SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering
Vipul Gupta, Zhuowan Li, Adam Kortylewski, Chenyu Zhang, Yingwei Li, Alan L. Yuille
Abstract
While Visual Question Answering (VQA) has progressed rapidly, previous works raise concerns about robustness of current VQA models. In this work, we study the robustness of VQA models from a novel perspective: visual context. We suggest that the models over-rely on the visual context, i.e., irrelevant objects in the image, to make predictions. To diagnose the models' reliance on visual context and measure their robustness, we propose a simple yet effective perturbation technique, SwapMix. SwapMix perturbs the visual context by swapping features of irrelevant context objects with features from other objects in the dataset. Using SwapMix we are able to change answers to more than 45% of the questions for a representative VQA model. Additionally, we train the models with perfect sight and find that the context over-reliance highly depends on the quality of visual representations. In addition to diagnosing, SwapMix can also be applied as a data augmentation strategy during training in order to regularize the context over-reliance. By swapping the context object features, the model reliance on context can be suppressed effectively. Two representative VQA models are studied using SwapMix: a co-attention model MCAN and a large-scale pretrained model LXMERT. Our experiments on the popular GQA dataset show the effectiveness of SwapMix for both diagnosing model robustness, and regularizing the over-reliance on visual context. The code for our method is available at https://github.com/vipulgupta1011/swapmix
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03d18119-1775-4111-b72e-b7f7dc17f8d1Cited by top-tier papers20
- Sim VQA: Exploring Simulated Environments for Visual Question AnsweringPaola Cascante-Bonilla, Hui Wu, Letao Wang, Rogério Feris et al.CVPR 2022 · 28 citations
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski et al.NeurIPS 2023 · 27 citations
- VisFIS: Visual Feature Importance Supervision with Right-for-the-Right-Reason ObjectivesZhuofan Ying, Peter Hase, Mohit BansalNeurIPS 2022 · 16 citations
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?Hao Yin, Guangzong Si, Zilei WangNeurIPS 2025 · 6 citations
- ReactioNet: Learning High-order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation ReasoningXiaotian Li, Taoyue Wang, Geran Zhao, Xiang Zhang et al.ICCV 2023 · 3 citations
Builds on10
- Taking a HINT: Leveraging Explanations to Make Vision and Language Models More GroundedRamprasaath Ramasamy Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin et al.ICCV 2019 · 288 citations
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 136 citations
- Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question AnsweringCorentin Dancette, Rémi Cadène, Damien Teney, Matthieu CordICCV 2021 · 95 citations
- Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial OcclusionAdam Kortylewski, Ju He, Qing Liu, Alan L. YuilleCVPR 2020
- Counterfactual Samples Synthesizing for Robust Visual Question AnsweringLong Chen, Xin Yan, Jun Xiao, Hanwang Zhang et al.CVPR 2020
Related papers
- Towards Causal VQA: Revealing and Reducing Spurious Correlations by Invariant and Covariant Semantic EditingVedika Agarwal, Rakshith Shetty, Mario FritzCVPR 2020
- Is Your (Reasoning) Multimodal Language Model Vulnerable Toward Distractions?Ming Liu, Hao Chen, Jindong Wang, Liwen Wang et al.AAAI 2026
- VQAttack: Transferable Adversarial Attacks on Visual Question Answering via Pre-trained ModelsZiyi Yin, Muchao Ye, Tianrong Zhang, Jiaqi Wang et al.AAAI 2024 · 20 citations
- Medical Vision-Language Pre-training with Multimodal Variational Masked Autoencoder for Robust Medical VQADexuan Xu, Yanyuan Chen, Yu Huang, Shihao E et al.ACM MM 2025
- Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich DocumentsDavide Napolitano, Luca Cagliero, Fabrizio BattiloroAAAI 2026
