Show, Deconfound and Tell: Image Captioning with Causal Inference
Bing Liu, Dong Wang, Xu Yang, Yong Zhou, Rui Yao, Zhiwen Shao, Jiaqi Zhao
摘要
The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https: //github.com/CUMTGG/CIIC .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain AdaptationZhengfeng Lai, Noranart Vesdapunt, Ning Zhou, Jun Wu 等ICCV 2023 · 被引用 90 次
- Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door AdjustmentCongzhi Zhang, Linhai Zhang, Deyu ZhouAAAI 2024 · 被引用 32 次
- Vision-and-Language Navigation via Causal LearningLiuyi Wang, Zongtao He, Ronghao Dang, Mengjiao Shen 等CVPR 2024 · 被引用 24 次
- Variational Causal Inference Network for Explanatory Visual Question AnsweringDizhan Xue, Shengsheng Qian, Changsheng XuICCV 2023 · 被引用 19 次
- Causal Representation Learning via Counterfactual InterventionXiutian Li, Siqi Sun, Rui FengAAAI 2024 · 被引用 13 次
它引用的顶会 Paper13
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao 等AAAI 2021 · 被引用 349 次
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 被引用 346 次
- Interventional Few-Shot LearningZhongqi Yue, Hanwang Zhang, Qianru Sun, Xian-Sheng HuaNeurIPS 2020 · 被引用 284 次
- Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer NetworkJiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen 等AAAI 2021 · 被引用 206 次
相关 Paper
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang 等ACM MM 2023 · 被引用 12 次
- Causal Attention for Vision-Language TasksXu Yang, Hanwang Zhang, Guojun Qi, Jianfei CaiCVPR 2021
- DeVLBert: Learning Deconfounded Visio-Linguistic RepresentationsShengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang 等ACM MM 2020 · 被引用 66 次
- Contextual Debiasing for Visual Recognition with Causal MechanismsRuyang Liu, Hao Liu, Ge Li, Haodi Hou 等CVPR 2022 · 被引用 42 次
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
