Towards Deconfounded Image-Text Matching with Causal Inference
Wenhui Li, Xinqi Su, Dan Song, Lanjun Wang, Kun Zhang, An-An Liu
摘要
Prior image-text matching methods have shown remarkable performance on many benchmark datasets, but most of them overlook the bias in the dataset, which exists in intra-modal and inter-modal, and tend to learn the spurious correlations that extremely degrade the generalization ability of the model. Furthermore, these methods often incorporate biased external knowledge from large-scale datasets as prior knowledge into image-text matching model, which is inevitable to force model further learn biased associations. To address above limitations, this paper firstly utilizes Structural Causal Models (SCMs) to illustrate how intra- and inter-modal confounders damage the image-text matching. Then, we employ backdoor adjustment to propose an innovative Deconfounded Causal Inference Network (DCIN) for image-text matching task. DCIN (1) decomposes the intra- and inter-modal confounders and incorporates them into the encoding stage of visual and textual features, effectively eliminating the spurious correlations during image-text matching, and (2) uses causal inference to mitigate biases of external knowledge. Consequently, the model can learn causality instead of spurious correlations caused by dataset bias. Extensive experiments on two well-known benchmark datasets, i.e., Flickr30K and MSCOCO, demonstrate the superiority of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- CWNet: Causal Wavelet Network for Low-Light Image EnhancementTongshun Zhang, Pingping Liu, Yubing Lu, Mengen Cai 等ICCV 2025 · 被引用 17 次
- INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zhiheng Fu, Zixu Li 等AAAI 2026 · 被引用 12 次
- Fair Deepfake Detectors Can GeneralizeHarry Cheng, Ming-Hui Liu, Yangyang Guo, Tianyi Wang 等NeurIPS 2025 · 被引用 11 次
- Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken GenerationYongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu 等SIGIR 2025 · 被引用 6 次
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li 等AAAI 2026
它引用的顶会 Paper19
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng 等ICCV 2019 · 被引用 349 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie 等ACM MM 2020 · 被引用 159 次
相关 Paper
- Show, Deconfound and Tell: Image Captioning with Causal InferenceBing Liu, Dong Wang, Xu Yang, Yong Zhou 等CVPR 2022 · 被引用 66 次
- Towards Unbiased Visual Emotion Recognition via Causal InterventionYuedong Chen, Xu Yang, Tat-Jen Cham, Jianfei CaiACM MM 2022 · 被引用 27 次
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu 等CVPR 2021
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- CausalCtrl: Causality-Aware Control Framework for Text-Guided Visual EditingHaoxiang Cao, Chaoqun Wang, Yongwen Lai, Shaobo Min 等ACM MM 2025 · 被引用 1 次
