Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
Yue Tan, Xiaoqian Hu, Hao Xue, Celso de Melo, Flora D. Salim
Abstract
Frontier vision-language models (VLMs) have made remarkable improvements in video understanding tasks. However, real-world videos typically exist as continuously evolving data streams (e.g., dynamic scenes captured by wearable glasses), necessitating models to continually adapt to shifting data distributions and novel scenarios. Considering the prohibitive computational costs of fine-tuning models on new tasks, usually, a small subset of parameters is updated while the bulk of the model remains frozen. This poses new challenges to existing continual learning frameworks in the context of large multimodal foundation models, i.e., catastrophic forgetting and update conflict. While the foundation models struggle with parameter-efficient continual learning, the hippocampus in the human brain has evolved highly efficient mechanisms for memory formation and consolidation. Inspired by the rapid Binding and pattern separation mechanisms in the hippocampus, in this work, we propose Bisecle for video-language continual learning, where a multi-directional supervision module is used to capture more cross-modal relationships and a contrastive prompt learning scheme is designed to isolate task-specific knowledge to facilitate efficient memory storage. Binding and separation processes further strengthen the ability of VLMs to retain complex experiences, enabling robust and efficient continual learning in video understanding tasks. We perform a thorough evaluation of the proposed Bisecle, demonstrating its ability to mitigate forgetting and enhance cross-task generalization on several VideoQA benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74f0af68-018e-48fb-b9f7-857b95c41c6dCited by top-tier papers8
- Assemble Your Crew: Automatic Multi-agent Communication Topology Design via Autoregressive Graph GenerationShiyuan Li, Yixin Liu, Qingsong Wen, Chengqi Zhang et al.AAAI 2026 · 29 citations
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen et al.CVPR 2026 · 7 citations
- Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly DetectionJunjun Pan, Yixin Liu, Rui Miao, Kaize Ding et al.ACL 2026 · 6 citations
- Correcting False Alarms from Unseen: Adapting Graph Anomaly Detectors at Test TimeJunjun Pan, Yixin Liu, Chuan Zhou, Fei Xiong et al.AAAI 2026 · 5 citations
- Towards One-for-All Anomaly Detection for Tabular DataShiyuan Li, Yixin Liu, Yu Zheng, Xiaofeng Cao et al.ICML 2026 · 3 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- ERNIE 2.0: A Continual Pre-Training Framework for Language UnderstandingYu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng et al.AAAI 2020 · 885 citations
Related papers
- BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language ModelingYizhao Gao, Nanyi Fei, Haoyu Lu, Zhiwu Lu et al.NeurIPS 2022 · 4 citations
- LADA: Scalable Label-Specific CLIP Adapter for Continual LearningMao-Lin Luo, Zi-Hao Zhou, Tong Wei, Min-Ling ZhangICML 2025
- HypCL: Adapting CLIP in Hyperbolic Space for Continual LearningQuan Cheng, Hao Yu, Da-Wei Zhou, Lijun ZhangICML 2026
- Fed-Duet: Dual Expert-Orchestrated Framework for Continual Federated Vision-Language LearningTao Guo, Junwei Chen, Laizhong CuiICLR 2026
- Calibrating Prompt from History for Continual Vision-Language Retrieval and GroundingTao Jin, Weicai Yan, Ye Wang, Sihang Cai et al.ACM MM 2024 · 6 citations
