X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval
Satya Krishna Gorti, Noël Vouitsis, Junwei Ma, Keyvan Golestan, Maksims Volkovs, Animesh Garg, Guangwei Yu
摘要
In text-video retrieval, the objective is to learn a cross-modal similarity function between a text and a video that ranks relevant text-video pairs higher than irrelevant pairs. However, videos inherently express a much wider gamut of information than texts. Instead, texts often capture sub-regions of entire videos and are most semantically similar to certain frames within videos. Therefore, for a given text, a retrieval model should focus on the text's most semantically similar video sub-regions to make a more relevant comparison. Yet, most existing works aggregate entire videos with-out directly considering text. Common text-agnostic ag-gregations schemes include mean-pooling or self-attention over the frames, but these are likely to encode misleading vi-sual information not described in the given text. To address this, we propose a cross-modal attention model called X-Pool that reasons between a text and the frames of a video. Our core mechanism is a scaled dot product attention for a text to attend to its most semantically similar frames. We then generate an aggregated video representation conditioned on the text's attention weights over the frames. We evaluate our method on three benchmark datasets of MSR-VTT, MSVD and LSMDC, achieving new state-of-the-art re-sults by up to 12% in relative improvement in Recall@ 1. Our findings thereby highlight the importance of joint text-video reasoning to extract important visual cues according to text. Full code and demo can be found at: layer6ai-labs.github.iolxpooll.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper85
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Expectation-Maximization Contrastive Learning for Compact Video-and-Language RepresentationsPeng Jin, Jinfa Huang, Fenglin Liu, Xian Wu 等NeurIPS 2022 · 被引用 105 次
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ICCV 2023 · 被引用 98 次
- DiffusionRet: Generative Text-Video Retrieval with Diffusion ModelPeng Jin, Hao Li, Zesen Cheng, Kehan Li 等ICCV 2023 · 被引用 95 次
- Unified Coarse-to-Fine Alignment for Video-Text RetrievalZiyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius 等ICCV 2023 · 被引用 90 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training ParadigmYangguang Li, Feng Liang, Lichen Zhao, Yufeng Cui 等ICLR 2022 · 被引用 565 次
相关 Paper
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan 等ACM MM 2022 · 被引用 314 次
- DGL: Dynamic Global-Local Prompt Tuning for Text-Video RetrievalXiangpeng Yang, Linchao Zhu, Xiaohan Wang, Yi YangAAAI 2024 · 被引用 53 次
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 被引用 7 次
- T2VLAD: Global-Local Sequence Alignment for Text-Video RetrievalXiaohan Wang, Linchao Zhu, Yi YangCVPR 2021
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen 等ICCV 2023 · 被引用 52 次
