Continual Text-to-Video Retrieval with Frame Fusion and Task-Aware Routing
Zecheng Zhao, Zhi Chen, Zi Huang, Shazia Sadiq, Tong Chen
Abstract
Text-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries. However, as video content evolves continuously, adapting TVR systems to new data remains a critical yet underexplored challenge. In this paper, we introduce the first benchmark for Continual Text-to-Video Retrieval (CTVR) to address the limitations of existing approaches. Current Pre-Trained Model (PTM)based TVR methods struggle with maintaining model plasticity when adapting to new tasks, while existing Continual Learning (CL) methods suffer from catastrophic forgetting, leading to semantic misalignment between historical queries and stored video features. To address these two challenges, we propose FrameFu-sionMoE, a novel CTVR framework that comprises two key components: (1) the Frame Fusion Adapter (FFA), which captures temporal video dynamics while preserving model plasticity, and (2) the Task-Aware Mixture-of-Experts (TAME), which ensures consistent semantic alignment between queries across tasks and the stored video features. Thus, FrameFusionMoE enables effective adaptation to new video content while preserving historical textvideo relevance to mitigate catastrophic forgetting. We comprehensively evaluate FrameFusionMoE on two benchmark datasets under various task settings. Results demonstrate that FrameFusionMoE outperforms existing CL and TVR methods, achieving superior retrieval performance with minimal degradation on earlier tasks when handling continuous video streams. Our code is available at: https://github.com/JasonCodeMaker/CTVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a14662b0-b0df-4f12-939f-a91085d9d9d0Cited by top-tier papers4
- SVIP: Semantically Contextualized Visual Patches for Zero-Shot LearningZhi Chen, Zecheng Zhao, Jingcai Guo, Jingjing Li et al.ICCV 2025 · 8 citations
- Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty MinimizationBingqing Zhang, Zhuo Cao, Heming Du, Yang Li et al.ICCV 2025 · 3 citations
- ProEx: A Unified Framework Leveraging Large Language Model with Profile Extrapolation for RecommendationYi Zhang, Yiwen Zhang, Yu Wang, Tong Chen et al.KDD 2026 · 1 citation
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu et al.SIGIR 2026
Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
Related papers
- Continual Predictive Learning from VideosGeng Chen, Wendong Zhang, Han Lu, Siyu Gao et al.CVPR 2022 · 5 citations
- BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language ModelingYizhao Gao, Nanyi Fei, Haoyu Lu, Zhiwu Lu et al.NeurIPS 2022 · 4 citations
- Bring Your Dreams to Life: Continual Text-to-Video CustomizationJiahua Dong, Xudong Wang, Wenqi Liang, Zongyan Han et al.AAAI 2026 · 1 citation
- vCLIMB: A Novel Video Class Incremental Learning BenchmarkAndrés Villa, Kumail Alhamoud, Victor Escorcia, Fabian Caba Heilbron et al.CVPR 2022 · 32 citations
- Prompt Switch: Efficient CLIP Adaptation for Text-Video RetrievalChaorui Deng, Qi Chen, Pengda Qin, Da Chen et al.ICCV 2023 · 52 citations
