Dual Learning with Dynamic Knowledge Distillation for Partially Relevant Video Retrieval
Jianfeng Dong, Minsong Zhang, Zheng Zhang, Xianke Chen, Daizong Liu, Xiaoye Qu, Xun Wang, Baolong Liu
Abstract
Almost all previous text-to-video retrieval works assume that videos are pre-trimmed with short durations. However, in practice, videos are generally untrimmed containing much background content. In this work, we investigate the more practical but challenging Partially Relevant Video Retrieval (PRVR) task, which aims to retrieve partially relevant untrimmed videos with the query input. Particularly, we propose to address PRVR from a new perspective, i.e., distilling the generalization knowledge from the large-scale vision-language pre-trained model and transferring it to a task-specific PRVR network. To be specific, we introduce a Dual Learning framework with Dynamic Knowledge Distillation (DL-DKD), which exploits the knowledge of a large vision-language model as the teacher to guide a student model. During the knowledge distillation, an inheritance student branch is devised to absorb the knowledge from the teacher model. Considering that the large model may be of mediocre performance due to the domain gaps, we further develop an exploration student branch to take the benefits of task-specific information. In addition, a dynamical knowledge distillation strategy is further devised to adjust the effect of each student branch learning during the training. Experiment results demonstrate that our proposed model achieves state-of-the-art performance on ActivityNet and TVR datasets for PRVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f95b628-9f82-438b-bea1-2f250697a773Cited by top-tier papers15
- Partially Relevant Video RetrievalJianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang et al.ACM MM 2022 · 65 citations
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou et al.AAAI 2024 · 30 citations
- CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge TransferYabing Wang, Fan Wang, Jianfeng Dong, Hao LuoAAAI 2024 · 20 citations
- Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual RetrievalZhe Ma, Jianfeng Dong, Shouling Ji, Zhenguang Liu et al.AAAI 2024 · 14 citations
- Temporal Sentence Grounding with Relevance Feedback in VideosJianfeng Dong, Xiaoman Peng, Daizong Liu, Xiaoye Qu et al.NeurIPS 2024 · 12 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Vision-Language Pre-Training with Triple Contrastive LearningJinyu Yang, Jiali Duan, Son Tran, Yi Xu et al.CVPR 2022 · 266 citations
Related papers
- Action-and-object Aware Alignment for Partially Relevant Video RetrievalChuanshen Chen, Kai Zhou, Zhiquan Wen, Zeng You et al.AAAI 2026
- Multi-Modal Inductive Framework for Text-Video RetrievalQian Li, Yucheng Zhou, Cheng Ji, Feihong Lu et al.ACM MM 2024 · 8 citations
- Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment RetrievalYujia Zhu, Hao Yang, Yibo Zhao, Chunjie Ma et al.ACM MM 2025
- Filling the Information Gap between Video and Query for Language-Driven Moment RetrievalDaizong Liu, Xiaoye Qu, Jianfeng Dong, Guoshun Nan et al.ACM MM 2023 · 10 citations
- Improving Video Model Transfer with Dynamic Representation LearningYi Li, Nuno VasconcelosCVPR 2022 · 2 citations
