CRET: Cross-Modal Retrieval Transformer for Efficient Text-Video Retrieval
Kaixiang Ji, Jiajia Liu, Weixiang Hong, Liheng Zhong, Jian Wang, Jingdong Chen, Wei Chu
Abstract
Given a text query, the text-to-video retrieval task aims to find the relevant videos in the database. Recently, model-based (MDB) methods have demonstrated superior accuracy than embedding-based (EDB) methods due to their excellent capacity of modeling local video/text correspondences, especially when equipped with large-scale pre-training schemes like ClipBERT. Generally speaking, MDB methods take a text-video pair as input and harness deep models to predict the mutual similarity, while EDB methods first utilize modality-specific encoders to extract embeddings for text and video, then evaluate the distance based on the extracted embeddings. Notably, MDB methods cannot produce explicit representations for text and video, instead, they have to exhaustively pair the query with every database item to predict their mutual similarities in the inference stage, which results in significant inefficiency in practical applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 791a5ce2-6511-4221-9e6c-2fdaed1f2531Cited by top-tier papers4
- Progressive Spatio-Temporal Prototype Matching for Text-Video RetrievalPandeng Li, Chen-Wei Xie, Liming Zhao, Hongtao Xie et al.ICCV 2023 · 62 citations
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng et al.NeurIPS 2025 · 6 citations
- Learning Dynamic Similarity by Bidirectional Hierarchical Sliding Semantic Probe for Efficient Text Video RetrievalYang Liu, Shudong Huang, Deng Xiong, Jiancheng LvAAAI 2025 · 4 citations
- Rethinking the Multimodal Correlation of Multimodal Sequential Learning via Generalizable Attentional Results AlignmentTao Jin, Wang Lin, Ye Wang, Linjun Li et al.ACL 2024 · 2 citations
Related papers
- Match4Match: Enhancing Text-Video Retrieval by Maximum Flow with Minimum CostZhongjie Duan, Chengyu Wang, Cen Chen, Wenmeng Zhou et al.WWW 2023 · 2 citations
- Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalArun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates et al.CVPR 2025
- MEME: Multi-Encoder Multi-Expert Framework with Data Augmentation for Video RetrievalSeong-Min Kang, Yoon-Sik ChoSIGIR 2023 · 7 citations
- Tree-Augmented Cross-Modal Encoding for Complex-Query Video RetrievalXun Yang, Jianfeng Dong, Yixin Cao, Xun Wang et al.SIGIR 2020 · 131 citations
- Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang et al.CVPR 2023
