Space-time Prompting for Video Class-incremental Learning
Yixuan Pei, Zhiwu Qing, Shiwei Zhang, Xiang Wang, Yingya Zhang, Deli Zhao, Xueming Qian
Abstract
Recently, prompt-based learning has made impressive progress on image class-incremental learning, but it still lacks sufficient exploration in the video domain. In this paper, we will fill this gap by learning multiple prompts based on a powerful image-language pre-trained model, i.e., CLIP, making it fit for video class-incremental learning (VCIL). For this purpose, we present a space-time prompting approach (ST-Prompt) which contains two kinds of prompts, i.e., task-specific prompts and task-agnostic prompts. The task-specific prompts are to address the catastrophic forgetting problem by learning multi-grained prompts, i.e., spatial prompts, temporal prompts and comprehensive prompts, for accurate task identification. The task-agnostic prompts maintain a globally-shared prompt pool, which can empower the pre-trained image models with temporal perception abilities by exchanging contexts between frames. By this means, ST-Prompt can transfer the plentiful knowledge in the image-language pre-trained models to the VCIL task with only a tiny set of prompts to be optimized. To evaluate ST-Prompt, we conduct extensive experiments on three standard benchmarks. The results show that ST-Prompt can significantly surpass the state-of-the-art VCIL methods, especially it gains 9.06% on HMDB51 dataset under the 1 × 25 stage setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfccb456-20cd-4cd8-bd3b-9fc27246d213Cited by top-tier papers7
- StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental LearningHuaijie Wang, De Cheng, Guozhang Li, Zhipeng Xu et al.ICLR 2026 · 8 citations
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningJongseo Lee, Kyungho Bae, Kyle Min, Gyeong-Moon Park et al.ICCV 2025 · 2 citations
- Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningXueyi Zhang, Chengwei Zhang, Zheng Li, Xiyu Wang et al.AAAI 2026 · 1 citation
- CRAM: Large-Scale Video Continual Learning with Bootstrapped CompressionShivani Mall, João F. HenriquesICCV 2025
- CEL: Continual Ego, Exo, and Ego-Exo LearningHongwei Yan, Kanglei Zhou, Yuchen Liu, Qingyu Shi et al.ICML 2026
Builds on31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
Related papers
- STOP: Integrated Spatial-Temporal Dynamic Prompting for Video UnderstandingZichen Liu, Kunlun Xu, Bing Su, Xu Zou et al.CVPR 2025
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationJiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao et al.ICCV 2025
- Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental LearningXiaohan Zou, Wenchao Ma, Shu ZhaoCVPR 2025
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang et al.AAAI 2024 · 54 citations
- PrePrompt: Predictive Prompting for Class Incremental LearningLibo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An et al.KDD 2026 · 4 citations
