Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
Chan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang, Semin Myeong, Sang-hyo Park, Hyeyoung Park
Abstract
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to capture the rich semantics, including temporal changes, inherent in the video. In addition, incorrect information caused by generative models can lead to inaccurate retrieval. To address these issues, we propose a new framework, Narrating the Video (NarVid), which strategically leverages the comprehensive information available from frame-level captions, the narration. The proposed NarVid exploits narration in multiple ways: 1) feature enhancement through crossmodal interactions between narration and video, 2) queryaware adaptive filtering to suppress irrelevant or incorrect information, 3) dual-modal matching score by adding query-video similarity and query-narration similarity, and 4) hard-negative loss to learn discriminative features from multiple perspectives using the two similarities from different views. Experimental results demonstrate that NarVid achieves state-of-the-art performance on various benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element GuidanceHuy Le, Nhat Chung, Tung Kieu, Anh Nguyen et al.ACM MM 2025 · 2 citations
- Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal RetrievalZhengxin Pan, Haishuai Wang, Fangyu Wu, Peng Zhang et al.ACM MM 2025 · 2 citations
- Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalZhiqian Zhao, Liang Li, Lei Shen, Xichun Sheng et al.AAAI 2026 · 1 citation
- T2VParser: Adaptive Decomposition Tokens for Partial Alignment in Text to Video RetrievalYili Li, Gang Xiong, Gaopeng Gou, Xiangyan Qu et al.ACM MM 2025
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?Wenhao Wu, Haipeng Luo, Bo Fang, Jingdong Wang et al.CVPR 2023
- Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal InteractionMingda Jia, Weiliang Meng, Zenghuang Fu, Yiheng Li et al.AAAI 2026 · 1 citation
- Leveraging Weighted Cross-Graph Attention for Visual and Semantic Enhanced Video Captioning NetworkDeepali Verma, Arya Haldar, Tanima DuttaAAAI 2023 · 13 citations
- Generation-Augmented Video Corpus Moment RetrievalMingjin Kuai, Qianyin Xiao, Juncheng Li, Jin Peng et al.SIGIR 2026
- Composed Video Retrieval via Enriched Context and Discriminative EmbeddingsOmkar Thawakar, Muzammal Naseer, Rao Muhammad Anwer, Salman H. Khan et al.CVPR 2024 · 10 citations
