Text-Adaptive Multiple Visual Prototype Matching for Video-Text Retrieval
Chengzhi Lin, Ancong Wu, Junwei Liang, Jun Zhang, Wenhang Ge, Wei-Shi Zheng, Chunhua Shen
摘要
Cross-modal retrieval between videos and texts has gained increasing research interest due to the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part of the information. Thus, a video can correspond to multiple different text descriptions and queries. We call this phenomenon the ``Video-Text Correspondence Ambiguity'' problem. Current techniques mostly concentrate on mining local or multi-level alignment between contents of a video and text (e.g., object to entity and action to verb). It is difficult for these methods to alleviate the video-text correspondence ambiguity by describing a video using only one single feature, which is required to be matched with multiple different text features at the same time. To address this problem, we propose a Text-Adaptive Multiple Visual Prototype Matching model, which automatically captures multiple prototypes to describe a video by adaptive aggregation of video token features. Given a query text, the similarity is determined by the most similar prototype to find correspondence in the video, which is termed text-adaptive matching. To learn diverse prototypes for representing the rich information in videos, we propose a variance loss to encourage different prototypes to attend to different contents of the video. Our method outperforms state-of-the-art methods on four public video retrieval datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- UATVR: Uncertainty-Adaptive Text-Video RetrievalBo Fang, Wenhao Wu, Chang Liu, Yu Zhou 等ICCV 2023 · 被引用 98 次
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie 等ICCV 2023 · 被引用 62 次
- Text Is MASS: Modeling as Stochastic Embedding for Text-Video RetrievalJiamian Wang, Pichao Wang, Guohao Sun, Dongfang Liu 等CVPR 2024 · 被引用 52 次
- Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video RetrievalCheol-Ho Cho, WonJun Moon, Woojin Jun, Minseok Jung 等AAAI 2025 · 被引用 11 次
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
相关 Paper
- Multi-event Video-Text RetrievalGengyuan Zhang, Jisen Ren, Jindong Gu, Volker TrespICCV 2023 · 被引用 19 次
- Uncertainty-Aware Alignment Network for Cross-Domain Video-Text RetrievalXiaoshuai Hao, Wanqian ZhangNeurIPS 2023 · 被引用 26 次
- T2VLAD: Global-Local Sequence Alignment for Text-Video RetrievalXiaohan Wang, Linchao Zhu, Yi YangCVPR 2021
- Prototypes Are Balanced Units for Efficient and Effective Partially Relevant Video RetrievalWonJun Moon, Cheol-Ho Cho, Woojin Jun, Taeoh Kim 等ICCV 2025 · 被引用 3 次
- Adaptive Cross-Modal Prototypes for Cross-Domain Visual-Language RetrievalYang Liu, Qingchao Chen, Samuel AlbanieCVPR 2021
