Discrete-Continuous Action Space Policy Gradient-Based Attention for Image-Text Matching
Shiyang Yan, Li Yu, Yuan Xie
摘要
Image-text matching is an important multi-modal task with massive applications. It tries to match the image and the text with similar semantic information. Existing approaches do not explicitly transform the different modalities into a common space. Meanwhile, the attention mechanism which is widely used in image-text matching models does not have supervision. We propose a novel attention scheme which projects the image and text embedding into a common space and optimises the attention weights directly towards the evaluation metrics. The proposed attention scheme can be considered as a kind of supervised attention and requiring no additional annotations. It is trained via a novel Discrete-continuous action space policy gradient algorithm, which is more effective in modelling complex action space than previous continuous action space policy gradient. We evaluate the proposed methods on two widelyused benchmark datasets: Flickr30k and MS-COCO, outperforming the previous approaches by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- Show Your Faith: Cross-Modal Confidence-Aware Network for Image-Text MatchingHuatian Zhang, Zhendong Mao, Kun Zhang, Yongdong ZhangAAAI 2022 · 被引用 62 次
- Identification of Necessary Semantic Undertakers in the Causal View for Image-Text MatchingHuatian Zhang, Lei Zhang, Kun Zhang, Zhendong MaoAAAI 2024 · 被引用 12 次
- Towards Deconfounded Image-Text Matching with Causal InferenceWenhui Li, Xinqi Su, Dan Song, Lanjun Wang 等ACM MM 2023 · 被引用 12 次
- Multilateral Semantic Relations Modeling for Image Text RetrievalZheng Wang, Zhenwei Gao, Kangshuai Guo, Yang Yang 等CVPR 2023
它引用的顶会 Paper2
相关 Paper
- Multi-Modality Cross Attention Network for Image and Sentence MatchingXi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang 等CVPR 2020
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 被引用 86 次
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 被引用 71 次
- Giving Text More Imagination Space for Image-text MatchingXinfeng Dong, Longfei Han, Dingwen Zhang, Li Liu 等ACM MM 2023 · 被引用 9 次
- Conceptual and Syntactical Cross-modal Alignment with Cross-level Consistency for Image-Text MatchingPengpeng Zeng, Lianli Gao, Xinyu Lyu, Shuaiqi Jing 等ACM MM 2021 · 被引用 37 次
