Causal Attention for Vision-Language Tasks
Xu Yang, Hanwang Zhang, Guojun Qi, Jianfei Cai
Abstract
We present a novel attention mechanism: Causal Attention (CATT), to remove the ever-elusive confounding effect in existing attention-based vision-language models. This effect causes harmful bias that misleads the attention module to focus on the spurious correlations in training data, damaging the model generalization. As the confounder is unobserved in general, we use the front-door adjustment to realize the causal intervention, which does not require any knowledge on the confounder. Specifically, CATT is implemented as a combination of 1) In-Sample Attention (IS-ATT) and 2) Cross-Sample Attention (CS-ATT), where the latter forcibly brings other samples into every IS-ATT, mimicking the causal intervention. CATT abides by the Q-K-V convention and hence can replace any attention module such as top-down attention and self-attention in Transformers. CATT improves various popular attention-based vision-language models by considerable margins. In particular, we show that CATT has great potential in large-scale pre-training, e.g., it can promote the lighter LXMERT [57] , which uses fewer data and less computational power, comparable to the heavier UNITER [14] . Code is published in https://github.com/yangxuntu/lxmertcatt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e412669b-20ee-4be8-89ff-499036ea2316Cited by top-tier papers56
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- Causal Attention for Interpretable and Generalizable Graph ClassificationYongduo Sui, Xiang Wang, Jiancan Wu, Min Lin et al.KDD 2022 · 166 citations
- Causal Attention for Unbiased Visual RecognitionTan Wang, Chang Zhou, Qianru Sun, Hanwang ZhangICCV 2021 · 162 citations
- C-CAM: Causal CAM for Weakly Supervised Semantic Segmentation on Medical ImageZhang Chen, Zhiqiang Tian, Jihua Zhu, Ce Li et al.CVPR 2022 · 116 citations
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 74 citations
Builds on15
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke et al.ICLR 2020 · 371 citations
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 284 citations
Related papers
- Show, Deconfound and Tell: Image Captioning with Causal InferenceBing Liu, Dong Wang, Xu Yang, Yong Zhou et al.CVPR 2022 · 66 citations
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min et al.ACM MM 2025 · 1 citation
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang et al.ACM MM 2020 · 66 citations
- Causality Meets the Table: Debiasing LLMs for Faithful TableQA via Front-Door InterventionZhen Yang, Ziwei Du, Minghan Zhang, Wei Du et al.NeurIPS 2025 · 6 citations
- Vision-and-Language Navigation via Causal LearningLiuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen et al.CVPR 2024 · 24 citations
