Multi-Modal Inductive Framework for Text-Video Retrieval
Qian Li, Yucheng Zhou, Cheng Ji, Feihong Lu, Jianian Gong, Shangguang Wang, Jianxin Li
摘要
Text-video retrieval (TVR) identifies relevant videos based on textual queries. Existing methods are limited by their ability to understand and connect different modalities, resulting in increased difficulty in retrievals. In this paper, we propose a generation-based TVR paradigm facilitated by LLM distillation to better learn and capture deep retrieval knowledge for text-video retrieval, amidsting the rapid evolution of Large Language Models. Specifically, we first design the fine-tuning large vision-language model that leverages the knowledge learned from language models to enhance the alignment of semantic information between the text and video modalities. It also incorporates an inductive reasoning mechanism, which focuses on incorporating important temporal and spatial features into the video embeddings. We further design question prompt clustering to select the most important prompts, considering their contribution to improving retrieval performance. Experimental results show that our approach achieves excellent performance on two benchmark datasets compared to its competitors.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu 等SIGIR 2026
- Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level CaptionsChan Hur, Jeong-Hun Hong, Dong-hun Lee, Dabin Kang 等CVPR 2025
- Gravitation-Driven Semantic Alignment for Text Video RetrievalYi Yang, Zheng Wang, Xing Xu, Jingkuan Song 等CVPR 2026
相关 Paper
- Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video RetrievalJianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen 等ICCV 2023 · 被引用 35 次
- ViG-RAG: Video-aware Graph Retrieval-Augmented Generation via Temporal and Semantic Hybrid ReasoningZongsheng Cao, Anran Liu, Yangfan He, Jing Li 等AAAI 2026 · 被引用 1 次
- Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction-Reasoning SynergyHaijier Chen, Bo Xu, Shoujian Zhang, Haoze Liu 等ICLR 2026 · 被引用 6 次
- ViLL-E: Video LLM Embeddings for RetrievalRohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei, Sheng Liu 等ACL 2026
- Bridging Information Asymmetry in Text-video Retrieval: A Data-centric ApproachZechen Bai, Tianjun Xiao, Tong He, Pichao Wang 等ICLR 2025
