Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and Regularization
Yawen Cui, Li Liu, Zitong Yu, Guanjie Huang, Xiaopeng Hong
摘要
Audio-Visual Learning (AVL) aims at the audio-visual perception with both audio and vision modalities. AVL also suffers from data insufficiency in many applications as with other unimodal tasks. Concurrently, AVL often needs to continuously learn over time rather than all knowledge simultaneously. Considering the above two perspectives, our work mainly focuses on benchmarking the unexplored Few-Shot Audio-Visual Class-Incremental Learning (FS-AVCIL), i.e., continually perceiving novel categories described by a limited number of labeled examples with audio and visual modalities. Firstly, we provide the detailed task configuration together with a thorough analysis of the challenges in FS-AVCIL:
(1) how to efficiently learn and fuse multimodal information with limited labeled examples; and (2) how to alleviate catastrophic forgetting cross-modal semantic correlations with limited data. Then, we propose an efficient framework based on Vision Transformer to solve FS-AVCIL, containing two parts: temporal-residual prompting for audio-visual synergy adapter and temporal prompt regularization. Specifically, temporal-residual prompting is incorporated into the audio-visual adapter to efficiently finetune the pre-trained foundation model with limited data and capture audio-visual correlation by learning temporal-relevant prompts. Besides, we regularize temporal-relevant prompts to memorize previous knowledge by fully using the temporal knowledge from various perspectives. This framework is validated in audiovisual classification tasks under the FS-AVCIL scenario, and extensive experiments demonstrate its superior performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen 等NeurIPS 2021 · 被引用 884 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen 等NeurIPS 2021 · 被引用 404 次
- Few-Shot Class-Incremental Learning via Relation Knowledge DistillationSonglin Dong, Xiaopeng Hong, Xiaoyu Tao, Xinyuan Chang 等AAAI 2021 · 被引用 215 次
- A Unified Continual Learning Framework with General Parameter-Efficient TuningQiankun Gao, Chen Zhao, Yifan Sun, Teng Xi 等ICCV 2023 · 被引用 152 次
相关 Paper
- Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic SegmentationJingqiao Xiu, Mengze Li, Zongxin Yang, Wei Ji 等AAAI 2025 · 被引用 3 次
- DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental LearningLinpu He, Yanan Li, Bingze Li, Elvis Han Cui 等ACM MM 2025 · 被引用 2 次
- Class-Incremental Grouping Network for Continual Audio-Visual LearningShentong Mo, Weiguo Pian, Yapeng TianICCV 2023 · 被引用 34 次
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
- Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental LearningJiong Yin, Liang Li, Jiehua Zhang, Yuhan Gao 等ICCV 2025 · 被引用 3 次
