Counterfactual Vision and Language Learning
Ehsan Abbasnejad, Damien Teney, Amin Parvaneh, Javen Shi, Anton van den Hengel
Abstract
The ongoing success of visual question answering methods has been somewhat surprising given that, at its most general, the problem requires understanding the entire variety of both visual and language stimuli. It is particularly remarkable that this success has been achieved on the basis of comparatively small datasets, given the scale of the problem. One explanation is that this has been accomplished partly by exploiting bias in the datasets rather than developing deeper multi-modal reasoning. This fundamentally limits the generalization of the method, and thus its practical applicability. We propose a method that addresses this problem by introducing counterfactuals in the training. In doing so we leverage structural causal models for counterfactual evaluation to formulate alternatives, for instance, questions that could be asked of the same image set. We show that simulating plausible alternative training data through this process results in better generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec11547d-8b52-42ef-9443-9630feec9ad1Cited by top-tier papers40
- On the Value of Out-of-Distribution Testing: An Example of Goodhart's LawDamien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha et al.NeurIPS 2020 · 163 citations
- Counterfactual Data-Augmented Sequential RecommendationZhenlei Wang, Jingsen Zhang, Hongteng Xu, Xu Chen et al.SIGIR 2021 · 131 citations
- Family as a Third Space for AI Literacies: How do children and parents learn about AI together?Stefania Druga, Fee Lia Christoph, Amy J. KoCHI 2022 · 118 citations
- Active Learning by Feature MixingAmin Parvaneh, Ehsan Abbasnejad, Damien Teney, Reza Haffari et al.CVPR 2022 · 113 citations
- Cross-View Geo-Localization via Learning Disentangled Geometric Layout CorrespondenceXiaohan Zhang, Xingyu Li, Waqas Sultani, Yi Zhou et al.AAAI 2023 · 111 citations
Related papers
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu et al.CVPR 2021
- Mitigating Language Bias of LMMs in Social Intelligence Understanding with Virtual Counterfactual CalibrationPeng Chen, Xiao-Yu Guo, Yuan-Fang Li, Xiaowang Zhang et al.EMNLP 2024 · 1 citation
- CF-VLM: CounterFactual Vision-Language Fine-tuningJusheng Zhang, Kaitong Cai, Yijia Fan, Jian Wang et al.NeurIPS 2025 · 71 citations
- What If the TV was off? Examining Counterfactual Reasoning Abilities of Multi-modal Language ModelsLetian Zhang, Xiaotong Zhai, Zhongkai Zhao, Yongshuo Zong et al.CVPR 2024 · 9 citations
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng et al.CVPR 2020
