Multi-Modality Cross Attention Network for Image and Sentence Matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, Feng Wu
摘要
The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence words for the cross-modal matching task. Different from them, in this work, we propose a novel Multi-Modality Cross Attention (MMCA) Network for image and sentence matching by jointly modeling the intra-modality and inter-modality relationships of image regions and sentence words in a unified deep model. In the proposed MM-CA, we design a novel cross-attention mechanism, which is able to exploit not only the intra-modality relationship within each modality, but also the inter-modality relationship between image regions and sentence words to complement and enhance each other for image and sentence matching. Extensive experimental results on two standard benchmarks including Flickr30K and MS-COCO demonstrate that the proposed model performs favorably against state-of-the-art image and sentence matching methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper55
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao 等SIGIR 2021 · 被引用 187 次
- Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth EstimationHyunyoung Jung, Eunhyeok Park, Sungjoo YooICCV 2021 · 被引用 133 次
- AutoFed: Heterogeneity-Aware Federated Multimodal Learning for Robust Autonomous DrivingTianyue Zheng, Ang Li, Zhe Chen, Hongbo Wang 等MobiCom 2023 · 被引用 75 次
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu 等ICML 2024 · 被引用 64 次
它引用的顶会 Paper1
相关 Paper
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
- Saliency-Guided Attention Network for Image-Sentence MatchingZhong Ji, Haoran Wang, Jungong Han, Yanwei PangICCV 2019 · 被引用 96 次
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 被引用 62 次
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing 等ACM MM 2021 · 被引用 37 次
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 被引用 71 次
