CRET: Cross-Modal Retrieval Transformer for Efficient Text-Video Retrieval
Kaixiang Ji, Jiajia Liu, Weixiang Hong, Liheng Zhong, Jian Wang, Jingdong Chen, Wei Chu
摘要
Given a text query, the text-to-video retrieval task aims to find the relevant videos in the database. Recently, model-based (MDB) methods have demonstrated superior accuracy than embedding-based (EDB) methods due to their excellent capacity of modeling local video/text correspondences, especially when equipped with large-scale pre-training schemes like ClipBERT. Generally speaking, MDB methods take a text-video pair as input and harness deep models to predict the mutual similarity, while EDB methods first utilize modality-specific encoders to extract embeddings for text and video, then evaluate the distance based on the extracted embeddings. Notably, MDB methods cannot produce explicit representations for text and video, instead, they have to exhaustively pair the query with every database item to predict their mutual similarities in the inference stage, which results in significant inefficiency in practical applications.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie 等ICCV 2023 · 被引用 62 次
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng 等NeurIPS 2025 · 被引用 6 次
- Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video RetrievalYang Liu, Shudong Huang, Deng Xiong, Jiancheng LvAAAI 2025 · 被引用 4 次
- Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results AlignmentTao Jin, Wang Lin, Ye Wang, Linjun Li 等ACL 2024 · 被引用 2 次
相关 Paper
- Match4Match: Enhancing Text-Video Retrieval by Maximum Flow with Minimum CostZhongjie Duan, Chengyu Wang, Cen Chen, Wenmeng Zhou 等WWW 2023 · 被引用 2 次
- Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalArun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates 等CVPR 2025
- MEME: Multi-Encoder Multi-Expert Framework with Data Augmentation for Video RetrievalSeong-Min Kang, Yoon-Sik ChoSIGIR 2023 · 被引用 7 次
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang 等SIGIR 2020 · 被引用 131 次
- Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang 等CVPR 2023
