Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video Retrieval
Bangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun, Zhanhui Kang, Gang Yang, Xirong Li
Abstract
The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers from low efficiency. In this study, we explore a new Hybrid-Tower framework that can hybridize the advantages of the Two-Tower and Single-Tower framework, achieving high effectiveness and efficiency simultaneously. We propose a novel hybrid method, Fine-grained Pseudo-query Interaction and Generation for T2VR, ie, PIG, which includes a new pseudo-query generator designed to generate a pseudo-query for each video. This enables the video feature and the textual features of pseudo-query to interact in a fine-grained manner, similar to the Single-Tower approaches to hold high effectiveness, even before the real textual query is received. Simultaneously, our method introduces no additional storage or computational overhead compared to the Two-Tower framework during the inference stage, thus maintaining high efficiency. Extensive experiments on five commonly used text-video retrieval benchmarks demonstrate that our method achieves a significant improvement over the baseline, with an increase of in R@1. Furthermore, our method matches the efficiency of Two-Tower models while achieving near state-of-the-art performance, highlighting the advantages of the Hybrid-Tower framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 856aece0-55c3-4322-9c2b-5d62ee035a0fCited by top-tier papers4
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
- Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video RetrievalJun Li, Peifeng Lai, Xuhang Lou, Jinpeng Wang et al.ICML 2026
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu et al.SIGIR 2026
- SAVE: Speech-Aware Video Representation Learning for Video-Text RetrievalRuixiang Zhao, Zhihao Xu, Bangxiang Lan, Zijie Xin et al.CVPR 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text RetrievalYiwei Ma, Guohai Xu, Xiaoshuai Sun, Ming Yan et al.ACM MM 2022 · 314 citations
- Unmasked Teacher: Towards Training-Efficient Video Foundation ModelsKunchang Li, Yali Wang, Yizhuo Li, Yi Wang et al.ICCV 2023 · 266 citations
Related papers
- Video-ColBERT: Contextualized Late Interaction for Text-to-Video RetrievalArun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates et al.CVPR 2025
- Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation LearningKaibin Tian, Yanhua Cheng, Yi Liu, Xinglin Hou et al.AAAI 2024 · 19 citations
- Holistic Features are Almost Sufficient for Text-to-Video RetrievalKaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan et al.CVPR 2024 · 15 citations
- DPDV: Dual-Pathway and Dual-View Representation Learning for Bridging Information Asymmetry in Text-Video RetrievalZequn Xie, Xin Liu, Fangming Feng, Boyun Zhang et al.ACL 2026
- TempMe: Video Temporal Token Merging for Efficient Text-Video RetrievalLeqi Shen, Tianxiang Hao, Tao He, Sicheng Zhao et al.ICLR 2025
