Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and Regularization
Yawen Cui, Li Liu, Zitong Yu, Guanjie Huang, Xiaopeng Hong
Abstract
Audio-Visual Learning (AVL) aims at the audio-visual perception with both audio and vision modalities. AVL also suffers from data insufficiency in many applications as with other unimodal tasks. Concurrently, AVL often needs to continuously learn over time rather than all knowledge simultaneously. Considering the above two perspectives, our work mainly focuses on benchmarking the unexplored Few-Shot Audio-Visual Class-Incremental Learning (FS-AVCIL), i.e., continually perceiving novel categories described by a limited number of labeled examples with audio and visual modalities. Firstly, we provide the detailed task configuration together with a thorough analysis of the challenges in FS-AVCIL:
(1) how to efficiently learn and fuse multimodal information with limited labeled examples; and (2) how to alleviate catastrophic forgetting cross-modal semantic correlations with limited data. Then, we propose an efficient framework based on Vision Transformer to solve FS-AVCIL, containing two parts: temporal-residual prompting for audio-visual synergy adapter and temporal prompt regularization. Specifically, temporal-residual prompting is incorporated into the audio-visual adapter to efficiently finetune the pre-trained foundation model with limited data and capture audio-visual correlation by learning temporal-relevant prompts. Besides, we regularize temporal-relevant prompts to memorize previous knowledge by fully using the temporal knowledge from various perspectives. This framework is validated in audiovisual classification tasks under the FS-AVCIL scenario, and extensive experiments demonstrate its superior performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b482627-067e-4a62-bb77-5a3d5dbad932Builds on14
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- Few-Shot Class-Incremental Learning via Relation Knowledge DistillationSonglin Dong, Xiaopeng Hong, Xiaoyu Tao, Xinyuan Chang et al.AAAI 2021 · 215 citations
- A Unified Continual Learning Framework with General Parameter-Efficient TuningQiankun Gao, Chen Zhao, Yifan Sun, Teng Xi et al.ICCV 2023 · 152 citations
Related papers
- Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic SegmentationJingqiao Xiu, Mengze Li, Zongxin Yang, Wei Ji et al.AAAI 2025 · 3 citations
- DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental LearningLinpu He, Yanan Li, Bingze Li, Elvis Han Cui et al.ACM MM 2025 · 2 citations
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 34 citations
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 44 citations
- Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningJiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao et al.ICCV 2025 · 3 citations
