Causal Attention for Unbiased Visual Recognition
Tan Wang, Chang Zhou, Qianru Sun, Hanwang Zhang
Abstract
Attention module does not always help deep models learn causal features that are robust in any confounding context, e.g., a foreground object feature is invariant to different backgrounds. This is because the confounders trick the attention to capture spurious correlations that benefit the prediction when the training and testing data are IID (identical & independent distribution); while harm the prediction when the data are OOD (out-of-distribution). The sole fundamental solution to learn causal attention is by causal intervention, which requires additional annotations of the confounders, e.g., a "dog" model is learned within "grass+dog" and "road+dog" respectively, so the "grass" and "road" contexts will no longer confound the "dog" recognition. However, such annotation is not only prohibitively expensive, but also inherently problematic, as the confounders are elusive in nature. In this paper, we propose a causal attention module (CaaM) that self-annotates the confounders in unsupervised fashion. In particular, multiple CaaMs can be stacked and integrated in conventional attention CNN and self-attention Vision Transformer. In OOD settings, deep models with CaaM outperform those without it significantly; even in IID settings, the attention localization is also improved by CaaM, showing a great potential in applications that require robust visual saliency. Codes are available at https://github.com/ Wangt-CN/CaaM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c911e27b-242d-4ba3-99b0-8740aeeecd41Cited by top-tier papers49
- Discovering Invariant Rationales for Graph Neural NetworksYingxin Wu, Xiang Wang, An Zhang, Xiangnan He et al.ICLR 2022 · 313 citations
- Learning Invariant Graph Representations for Out-of-Distribution GeneralizationHaoyang Li, Ziwei Zhang, Xin Wang, Wenwu ZhuNeurIPS 2022 · 170 citations
- Causal Attention for Interpretable and Generalizable Graph ClassificationYongduo Sui, Xiang Wang, Jiancan Wu, Min Lin et al.KDD 2022 · 166 citations
- Invariant Grounding for Video Question AnsweringYicong Li, Xiang Wang, Junbin Xiao, Wei Ji et al.CVPR 2022 · 108 citations
- Self-Supervised Learning Disentangled Group Representation as FeatureTan Wang, Zhongqi Yue, Jianqiang Huang, Qianru Sun et al.NeurIPS 2021 · 78 citations
Builds on21
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun et al.ICCV 2021 · 733 citations
Related papers
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- Causality Compensated Attention for Contextual Biased Visual RecognitionRuyang Liu, Jingjia Huang, Thomas H. Li, Ge LiICLR 2023
- Improving Weakly Supervised Object Localization via Causal InterventionFeifei Shao, Yawei Luo, Li Zhang, Lu Ye et al.ACM MM 2021 · 24 citations
- In Pursuit of Causal Label Correlations for Multi-label Image RecognitionZhao-Min Chen, Xin Jin, Yisu Ge, Sixian ChanNeurIPS 2024 · 9 citations
- A Causal Debiasing Framework for Unsupervised Salient Object DetectionXiangru Lin, Ziyi Wu, Guanqi Chen, Guanbin Li et al.AAAI 2022 · 34 citations
