Show, Deconfound and Tell: Image Captioning with Causal Inference
Bing Liu, Dong Wang, Xu Yang, Yong Zhou, Rui Yao, Zhiwen Shao, Jiaqi Zhao
Abstract
The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https: //github.com/CUMTGG/CIIC .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05e485d1-654e-4d1d-a992-2b21d44e802fCited by top-tier papers21
- PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain AdaptationZhengfeng Lai, Noranart Vesdapunt, Ning Zhou, Jun Wu et al.ICCV 2023 · 90 citations
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Deyu ZhouAAAI 2024 · 32 citations
- Vision-and-Language Navigation via Causal LearningLiuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen et al.CVPR 2024 · 24 citations
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 19 citations
- Causal Representation Learning via Counterfactual InterventionXiutian Li, Siqi Sun, Rui FengAAAI 2024 · 13 citations
Builds on13
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 284 citations
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen et al.AAAI 2021 · 206 citations
Related papers
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang et al.ACM MM 2023 · 12 citations
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang et al.ACM MM 2020 · 66 citations
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou et al.CVPR 2022 · 42 citations
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu et al.CVPR 2021
