Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning
Yixuan Pei, Zhiwu Qing, Jun Cen, Xiang Wang, Shiwei Zhang, Yaxiong Wang, Mingqian Tang, Nong Sang, Xueming Qian
Abstract
Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning approach that learns to produce a condensed frame for each selected video. Specifically, FrameMaker is mainly composed of two crucial components: Frame Condensing and Instance-Specific Prompt. The former is to reduce the memory cost by preserving only one condensed frame instead of the whole video, while the latter aims to compensate the lost spatio-temporal details in the Frame Condensing stage. By this means, FrameMaker enables a remarkable reduction in memory but keep enough information that can be applied to following incremental tasks. Experimental results on multiple challenging benchmarks, i.e., HMDB51, UCF101 and Something-Something V2, demonstrate that FrameMaker can achieve better performance to recent advanced methods while consuming only 20% memory. Additionally, under the same memory consumption conditions, FrameMaker significantly outperforms existing state-of-the-arts by a convincing margin. * equal contribution 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b276f1a7-71b3-4287-8863-5a6e6cf4d846Cited by top-tier papers5
- Hypercorrelation Evolution for Video Class-Incremental LearningSen Liang, Kai Zhu, Wei Zhai, Zhiheng Liu et al.AAAI 2024 · 4 citations
- Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningXueyi Zhang, Chengwei Zhang, Zheng Li, Xiyu Wang et al.AAAI 2026 · 1 citation
- CRAM: Large-Scale Video Continual Learning with Bootstrapped CompressionShivani Mall, João F. HenriquesICCV 2025
- Coherent Temporal Synthesis for Incremental Action SegmentationGuodong Ding, Hans Golong, Angela YaoCVPR 2024
- Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental LearningXiaohan Zou, Wenchao Ma, Shu ZhaoCVPR 2025
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- When Video Classification Meets Incremental ClassesHanbin Zhao, Xin Qin, Shihao Su, Yongjian Fu et al.ACM MM 2021 · 26 citations
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningJongseo Lee, Kyungho Bae, Kyle Min, Gyeong-Moon Park et al.ICCV 2025 · 2 citations
- Class-Incremental Learning for Action Recognition in VideosJaeyoo Park, Minsoo Kang, Bohyung HanICCV 2021 · 68 citations
- OCSampler: Compressing Videos to One Clip with Single-step SamplingJintao Lin, Haodong Duan, Kai Chen, Dahua Lin et al.CVPR 2022 · 27 citations
- TS-ILM: Class Incremental Learning for Online Action DetectionXiaochen Li, Jian Cheng, Ziying Xia, Zichong Chen et al.ACM MM 2024 · 2 citations
