Affordance-First Decomposition for Continual Learning in Video–Language Understanding
Mengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen, Canran Xiao
Abstract
Continual learning for video--language understanding is increasingly important as models face non-stationary data, domains, and query styles, yet prevailing solutions blur what should stay stable versus what should adapt, rely on static routing/capacity, or require replaying past videos. We aim to explicitly specify where stability lives and where plasticity should be focused under realistic memory and privacy constraints. We introduce Affordance-First Decomposition (AFD): videos are mapped to slowly varying affordance tokens that form a shared, time-aligned substrate, while a lightweight, query-routed, conflict-aware scheduler concentrates adaptation and grows capacity only when needed. The substrate is stabilized via weak alignment and teacher consistency, and training uses question-only replay. AFD achieves state-of-the-art across protocols: 51.6% average accuracy with -1.8% forgetting on domain-incremental VideoQA, ViLCo R@1@0.5 of 29.6% (MQ) and 20.7% (NLQ) with 18.4% stAP@0.25 (VQ), and 39.5% accuracy with -1.6% forgetting on time-incremental iVQA. Overall, AFD offers an explicit, interpretable split between a stable interaction-centered substrate and targeted adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ec069ef-2ba1-4bb7-a97d-c9904a5cf272Cited by top-tier papers4
- Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-TuningHua Ye, Siyuan Chen, Haoliang Zhang, Weihao Luo et al.NeurIPS 2025 · 2 citations
- Whose Instructions Count? Resolving Preference Bias in Instruction Fine-TuningJiayu Zhang, Changbang Li, Yinan Peng, Weihao Luo et al.NeurIPS 2025 · 2 citations
- Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal LearningJiayu Zhang, Chuangxin Zhao, Canran Xiao, Ruibo Duan et al.ICLR 2026
- CoMem: Compositional Concept-Graph Memory for Vision-Language AdaptationHeng Zhou, Jing Tang, Jusheng Zhang, Yanshu Li et al.ICLR 2026
Builds on29
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 397 citations
- A Unified Continual Learning Framework with General Parameter-Efficient TuningQiankun Gao, Chen Zhao, Yifan Sun, Teng Xi et al.ICCV 2023 · 152 citations
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin et al.ICCV 2023 · 133 citations
- Revisiting the "Video" in Video-Language UnderstandingShyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu et al.CVPR 2022 · 121 citations
Related papers
- Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringImad Eddine Marouf, Enzo Tartaglione, Stéphane Lathuilière, Joost van de WeijerICCV 2025 · 4 citations
- MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question AnsweringZhifei Li, Yiran Wang, Chenyi Xiong, Yujing Xia et al.AAAI 2026
- Reversible Primitive-Composition Alignment for Continual Vision-Language LearningCanran Xiao, Tianxiang Xu, SiYuan Ma, Yiyang Jiang et al.ICLR 2026
- Bridging the Grounding Gap in VideoQA via Typed Memory for Language-based Belief-State ReasoningSaman Forouzandeh, Wei Peng, Xinghuo Yu, Mahdi JaliliICML 2026
- VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory BridgesYuxuan Wang, Yiqi Song, Cihang Xie, Yang Liu et al.ICCV 2025 · 7 citations
