Identification of Necessary Semantic Undertakers in the Causal View for Image-Text Matching
Huatian Zhang, Lei Zhang, Kun Zhang, Zhendong Mao
Abstract
Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Finegrained semantic interactions come from fragment alignments between image regions and text words. However, not all fragments contribute to image-text relevance, and many existing methods are devoted to mining the vital ones to measure the relevance accurately. How well image and text relate depends on the degree of semantic sharing between them. Treating the degree as an effect and fragments as its possible causes, we define those indispensable causes for the generation of the degree as necessary undertakers, i.e., if any of them did not occur, the relevance would be no longer valid. In this paper, we revisit image-text matching in the causal view and uncover inherent causal properties of relevance generation. Then we propose a novel theoretical prototype for estimating the probability-of-necessity of fragments, PN f , for the degree of semantic sharing by means of causal inference, and further design a Necessary Undertaker Identification Framework (NUIF) for image-text matching, which explicitly formalizes the fragment's contribution to imagetext relevance by modeling PN f in two ways. Extensive experiments show that our method achieves state-of-the-art on benchmarks Flickr30K and MSCOCO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Homology Consistency Constrained Efficient Tuning for Vision-Language ModelsHuatian Zhang, Lei Zhang, Yongdong Zhang, Zhendong MaoNeurIPS 2024 · 5 citations
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 4 citations
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li et al.AAAI 2026
- DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningKun Zhang, Jingyu Li, Zhe Li, S. Kevin ZhouCVPR 2025
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng et al.CVPR 2026
Builds on28
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu et al.ICLR 2022 · 827 citations
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang et al.SIGIR 2021 · 198 citations
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang et al.CVPR 2022 · 190 citations
Related papers
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 185 citations
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 62 citations
- Synthesizing Counterfactual Samples for Effective Image-Text MatchingHao Wei, Shuhui Wang, Xinzhe Han, Zhe Xue et al.ACM MM 2022 · 11 citations
- Progressive Positive Association Framework for Image and Text RetrievalWenhui Li, Yan Wang, Yuting Su, Lanjun Wang et al.ACM MM 2023 · 4 citations
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
