Space-time Prompting for Video Class-incremental Learning
Yixuan Pei, Zhiwu Qing, Shiwei Zhang, Xiang Wang, Yingya Zhang, Deli Zhao, Xueming Qian
摘要
Recently, prompt-based learning has made impressive progress on image class-incremental learning, but it still lacks sufficient exploration in the video domain. In this paper, we will fill this gap by learning multiple prompts based on a powerful image-language pre-trained model, i.e., CLIP, making it fit for video class-incremental learning (VCIL). For this purpose, we present a space-time prompting approach (ST-Prompt) which contains two kinds of prompts, i.e., task-specific prompts and task-agnostic prompts. The task-specific prompts are to address the catastrophic forgetting problem by learning multi-grained prompts, i.e., spatial prompts, temporal prompts and comprehensive prompts, for accurate task identification. The task-agnostic prompts maintain a globally-shared prompt pool, which can empower the pre-trained image models with temporal perception abilities by exchanging contexts between frames. By this means, ST-Prompt can transfer the plentiful knowledge in the image-language pre-trained models to the VCIL task with only a tiny set of prompts to be optimized. To evaluate ST-Prompt, we conduct extensive experiments on three standard benchmarks. The results show that ST-Prompt can significantly surpass the state-of-the-art VCIL methods, especially it gains 9.06% on HMDB51 dataset under the 1 × 25 stage setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental LearningHuaijie Wang, De Cheng, Guozhang Li, Zhipeng Xu 等ICLR 2026 · 被引用 8 次
- ESSENTIAL: Episodic and Semantic Memory Integration for Video Class-Incremental LearningJongseo Lee, Kyungho Bae, Kyle Min, Gyeong-Moon Park 等ICCV 2025 · 被引用 2 次
- Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative LearningXueyi Zhang, Chengwei Zhang, Zheng Li, Xiyu Wang 等AAAI 2026 · 被引用 1 次
- CRAM: Large-Scale Video Continual Learning with Bootstrapped CompressionShivani Mall, João F. HenriquesICCV 2025
- CEL: Continual Ego, Exo, and Ego-Exo LearningHongwei Yan, Kanglei Zhou, Yuchen Liu, Qingyu Shi 等ICML 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
相关 Paper
- STOP: Integrated Spatial-Temporal Dynamic Prompting for Video UnderstandingZichen Liu, Kunlun Xu, Bing Su, Xu Zou 等CVPR 2025
- Hierarchical Visual Prompt Learning for Continual Video Instance SegmentationJiahua Dong, Hui Yin, Wenqi Liang, Hanbin Zhao 等ICCV 2025
- Learning Conditional Space-Time Prompt Distributions for Video Class-Incremental LearningXiaohan Zou, Wenchao Ma, Shu ZhaoCVPR 2025
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
- PrePrompt: Predictive Prompting for Class Incremental LearningLibo Huang, Xiangqi Li, Jiarui Zhao, Zhulin An 等KDD 2026 · 被引用 4 次
