Action-Aware Embedding Enhancement for Image-Text Retrieval
Jiangtong Li, Li Niu, Liqing Zhang
摘要
Image-text retrieval plays a central role in bridging vision and language, which aims to reduce the semantic discrepancy between images and texts. Most of existing works rely on refined words and objects representation through the data-oriented method to capture the word-object cooccurrence. Such approaches are prone to ignore the asymmetric action relation between images and texts, that is, the text has explicit action representation (i.e., verb phrase) while the image only contains implicit action information. In this paper, we propose Action-aware Memory-Enhanced embedding (AME) method for image-text retrieval, which aims to emphasize the action information when mapping the images and texts into a shared embedding space. Specifically, we integrate action prediction along with an action-aware memory bank to enrich the image and text features with action-similar text features. The effectiveness of our proposed AME method is verified by comprehensive experimental results on two benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Recipe for Scaling up Text-to-Video Generation with Text-free VideosXiang Wang, Shiwei Zhang, Hangjie Yuan, Zhiwu Qing 等CVPR 2024 · 被引用 19 次
- Knowledge Proxy Intervention for Deconfounded Video Question AnsweringJiangtong Li, Li Niu, Liqing ZhangICCV 2023 · 被引用 7 次
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 被引用 3 次
- Recurrent Feature Mining and Keypoint Mixup Padding for Category-Agnostic Pose EstimationJunjie Chen, Weilong Chen, Yifan Zuo, Yuming FangCVPR 2025
它引用的顶会 Paper10
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng 等ICCV 2019 · 被引用 349 次
- Context-Aware Multi-View Summarization Network for Image-Text MatchingLeigang Qu, Meng Liu, Da Cao, Liqiang Nie 等ACM MM 2020 · 被引用 159 次
- Adaptive Cross-Modal Embeddings for Image-Text AlignmentJonatas Wehrmann, Camila Kolling, Rodrigo C. BarrosAAAI 2020 · 被引用 86 次
- Expressing Objects Just Like Words: Recurrent Visual Embedding for Image-Text MatchingTianlang Chen, Jiebo LuoAAAI 2020 · 被引用 71 次
相关 Paper
- Structured Multi-modal Feature Embedding and Alignment for Image-Sentence RetrievalXuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji 等ACM MM 2021 · 被引用 49 次
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- IMRAM: Iterative Matching With Recurrent Attention Memory for Cross-Modal Image-Text RetrievalHui Chen, Guiguang Ding, Xudong Liu, Zijia Lin 等CVPR 2020
- Cross-Modal Coherence for Text-to-Image RetrievalMalihe Alikhani, Fangda Han, Hareesh Ravi, Mubbasir Kapadia 等AAAI 2022 · 被引用 11 次
- Webly Supervised Knowledge Embedding Model for Visual ReasoningWenbo Zheng, Lan Yan, Chao Gou, Fei-Yue WangCVPR 2020
