Identification of Necessary Semantic Undertakers in the Causal View for Image-Text Matching
Huatian Zhang, Lei Zhang, Kun Zhang, Zhendong Mao
摘要
Image-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Finegrained semantic interactions come from fragment alignments between image regions and text words. However, not all fragments contribute to image-text relevance, and many existing methods are devoted to mining the vital ones to measure the relevance accurately. How well image and text relate depends on the degree of semantic sharing between them. Treating the degree as an effect and fragments as its possible causes, we define those indispensable causes for the generation of the degree as necessary undertakers, i.e., if any of them did not occur, the relevance would be no longer valid. In this paper, we revisit image-text matching in the causal view and uncover inherent causal properties of relevance generation. Then we propose a novel theoretical prototype for estimating the probability-of-necessity of fragments, PN f , for the degree of semantic sharing by means of causal inference, and further design a Necessary Undertaker Identification Framework (NUIF) for image-text matching, which explicitly formalizes the fragment's contribution to imagetext relevance by modeling PN f in two ways. Extensive experiments show that our method achieves state-of-the-art on benchmarks Flickr30K and MSCOCO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Homology Consistency Constrained Efficient Tuning for Vision-Language ModelsHuatian Zhang, Lei Zhang, Yongdong Zhang, Zhendong MaoNeurIPS 2024 · 被引用 5 次
- Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingXiang Ma, Xuemei Li, Lexin Fang, Caiming ZhangACM MM 2024 · 被引用 4 次
- Explicit Modeling of Causal Factors and Confounders for Image ClassificationWei Wu, Lei Meng, Zhuang Qi, Zixuan Li 等AAAI 2026
- DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningKun Zhang, Jingyu Li, Zhe Li, S. Kevin ZhouCVPR 2025
- CoV-Align: Efficient Fine-grained Cross-Modal Alignment with Cohesive Visual Semantics PriorityHengqi Liu, Wanting Zhou, Longteng Kong, Fangxiang Feng 等CVPR 2026
它引用的顶会 Paper28
- FILIP: Fine-grained Interactive Language-Image Pre-TrainingLewei Yao, Runhui Huang, Lu Hou, Guansong Lu 等ICLR 2022 · 被引用 827 次
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua 等NeurIPS 2020 · 被引用 563 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- Deconfounded Video Moment Retrieval with Causal InterventionXun Yang, Fuli Feng, Wei Ji, Meng Wang 等SIGIR 2021 · 被引用 198 次
- Causality Inspired Representation Learning for Domain GeneralizationFangrui Lv, Jian Liang, Shuang Li, Bin Zang 等CVPR 2022 · 被引用 190 次
相关 Paper
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 被引用 62 次
- Synthesizing Counterfactual Samples for Effective Image-Text MatchingHao Wei, Shuhui Wang, Xinzhe Han, Zhe Xue 等ACM MM 2022 · 被引用 11 次
- Progressive Positive Association Framework for Image and Text RetrievalWenhui Li, Yan Wang, Yuting Su, Lanjun Wang 等ACM MM 2023 · 被引用 4 次
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
