Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text Matching
Huatian Zhang, Zhendong Mao, Kun Zhang, Yongdong Zhang
Abstract
Image-text matching bridges vision and language, which is a crucial task in the field of multi-modal intelligence. The key challenge lies in how to measure image-text relevance accurately as matching evidence. Most existing works aggregate the local semantic similarities of matched region-word pairs as the overall relevance, and they typically assume that the matched pairs are equally reliable. However, although a region-word pair is locally matched across modalities, it may be inconsistent/unreliable from the global perspective of image-text, resulting in inaccurate relevance measurement. In this paper, we propose a novel Cross-Modal Confidence-Aware Network to infer the matching confidence that indicates the reliability of matched region-word pairs, which is combined with the local semantic similarities to refine the relevance measurement. Specifically, we first calculate the matching confidence via the relevance between the semantic of image regions and the complete described semantic in the image, with the text as a bridge. Further, to richly express the region semantics, we extend the region to its visual context in the image. Then, local semantic similarities are weighted with the inferred confidence to filter out unreliable matched pairs in aggregating. Comprehensive experiments show that our method achieves state-of-the-art performance on benchmarks Flickr30K and MSCOCO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58b4b46c-b3ad-4f7f-9105-9a131bf25f95Cited by top-tier papers6
- Cross-modal Active Complementary Learning with Self-refining CorrespondenceYang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou et al.NeurIPS 2023 · 49 citations
- Identification of Necessary Semantic Undertakers in the Causal View for Image-Text MatchingHuatian Zhang, Lei Zhang, Kun Zhang, Zhendong MaoAAAI 2024 · 12 citations
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang et al.ACM MM 2023 · 12 citations
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu et al.EMNLP 2024 · 10 citations
- DH-Set: Improving Vision-Language Alignment with Diverse and Hybrid Set-Embeddings LearningKun Zhang, Jingyu Li, Zhe Li, S. Kevin ZhouCVPR 2025
Builds on11
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 86 citations
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 71 citations
Related papers
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
- Unlocking the Power of Cross-Dimensional Semantic Dependency for Image-Text MatchingKun Zhang, Lei Zhang, Bo Hu, Mengxiao Zhu et al.ACM MM 2023 · 19 citations
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie et al.ACM MM 2020 · 159 citations
