Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts
Bingqing Zhang, Zhuo Cao, Heming Du, Yang Li, Xue Li, Jiajun Liu, Sen Wang
Abstract
Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world query shifts, where the distribution of query data deviates from the training domain, leading to a sharp performance drop. Existing image-focused robustness solutions are inadequate to handle this vulnerability in video, as they fail to address the complex spatio-temporal dynamics inherent in these shifts. To systematically evaluate this vulnerability, we first introduce a comprehensive benchmark featuring 12 distinct types of video perturbations across five severity degrees. Analysis on this benchmark reveals that query shifts amplify the hubness phenomenon, where a few gallery items become dominant "hubs" that attract a disproportionate number of queries. To mitigate this, we then propose HAT-VTR (Hubness Alleviation for Test-time Video-Text Retrieval), as our baseline test-time adaptation framework designed to directly counteract hubness in VTR. It leverages two key components: a Hubness Suppression Memory to refine similarity scores, and multi-granular losses to enforce temporal feature consistency. Extensive experiments demonstrate that HAT-VTR substantially improves robustness, consistently outperforming prior methods across diverse query shift scenarios, and enhancing model reliability for real-world applications. Code is available at https://github.com/bingqingzhang/vtr_tta.git .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1d0e044-ffbf-4865-a7d2-6f1de02c35f7Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- Mind-the-Gap! Unsupervised Domain Adaptation for Text-Video RetrievalQingchao Chen, Yang Liu, Samuel AlbanieAAAI 2021 · 28 citations
- Adversarial Hubness in Multi-Modal RetrievalTingwei Zhang, Fnu Suya, Rishi D. Jha, Collin Zhang et al.S&P 2026 · 1 citation
- Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery BanksYimu Wang, Xiangru Jian, Bo XueEMNLP 2023 · 7 citations
- Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware RoutingZecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq et al.SIGIR 2025 · 6 citations
- Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal InconsistencyJiafeng Liang, Shixin Jiang, Xuan Dong, Ning Wang et al.ACL 2025 · 2 citations
