Multi-Modality Cross Attention Network for Image and Sentence Matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, Feng Wu
Abstract
The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence words for the cross-modal matching task. Different from them, in this work, we propose a novel Multi-Modality Cross Attention (MMCA) Network for image and sentence matching by jointly modeling the intra-modality and inter-modality relationships of image regions and sentence words in a unified deep model. In the proposed MM-CA, we design a novel cross-attention mechanism, which is able to exploit not only the intra-modality relationship within each modality, but also the inter-modality relationship between image regions and sentence words to complement and enhance each other for image and sentence matching. Extensive experimental results on two standard benchmarks including Flickr30K and MS-COCO demonstrate that the proposed model performs favorably against state-of-the-art image and sentence matching methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9394fa3-99f3-4304-a964-360ede41e626Cited by top-tier papers55
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
- Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth EstimationHyunyoung Jung, Eunhyeok Park, Sungjoo YooICCV 2021 · 133 citations
- AutoFed: Heterogeneity-Aware Federated Multimodal Learning for Robust Autonomous DrivingTianyue Zheng, Ang Li, Zhe Chen, Hongbo Wang et al.MobiCom 2023 · 75 citations
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
Builds on1
Related papers
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
- Saliency-Guided Attention Network for Image-Sentence MatchingZhong Ji, Haoran Wang, Jungong Han, Yanwei PangICCV 2019 · 96 citations
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 62 citations
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing et al.ACM MM 2021 · 37 citations
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 71 citations
