Saliency-Guided Attention Network for Image-Sentence Matching
Zhong Ji, Haoran Wang, Jungong Han, Yanwei Pang
Abstract
This paper studies the task of matching image and sentence, where learning appropriate representations to bridge the semantic gap between image contents and language appears to be the main challenge. Unlike previous approaches that predominantly deploy symmetrical architecture to represent both modalities, we introduce a Saliency-guided Attention Network (SAN) that is characterized by building an asymmetrical link between vision and language to efficiently learn a fine-grained cross-modal correlation. The proposed SAN mainly includes three components: saliency detector, Saliency-weighted Visual Attention (SVA) module, and Saliency-guided Textual Attention (STA) module. Concretely, the saliency detector provides the visual saliency information to drive both two attention modules. Taking advantage of the saliency information, SVA is able to learn more discriminative visual features. By fusing the visual information from SVA and intra-modal information as a multi-modal guidance, STA affords us powerful textual representations that are synchronized with visual clues. Extensive experiments demonstrate SAN can improve the state-of-the-art results on the benchmark Flickr30K and MSCOCO datasets by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 462b0e86-d236-476e-a4dc-adf513c06a3fCited by top-tier papers12
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 185 citations
- Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceSiyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee LimAAAI 2021 · 162 citations
- LapsCore: Language-guided Person Search via Color ReasoningYushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li et al.ICCV 2021 · 91 citations
- Disentangled High Quality Salient Object DetectionLv Tang, Bo Li, Yijie Zhong, Shouhong Ding et al.ICCV 2021 · 86 citations
Related papers
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang et al.CVPR 2020
- Graph Structured Network for Image-Text MatchingChunxiao Liu, Zhendong Mao, Tianzhu Zhang, Hongtao Xie et al.CVPR 2020
- Fine-grained Image-text Matching by Cross-modal Hard Aligning NetworkZhengxin Pan, Fangyu Wu, Bailing ZhangCVPR 2023
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 62 citations
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
