Class-Incremental Grouping Network for Continual Audio-Visual Learning
Shentong Mo, Weiguo Pian, Yapeng Tian
摘要
Continual learning is a challenging problem in which models need to be trained on non-stationary data across sequential tasks for class-incremental learning. While previous methods have focused on using either regularization or rehearsal-based frameworks to alleviate catastrophic forgetting in image classification, they are limited to a single modality and cannot learn compact class-aware cross-modal representations for continual audio-visual learning. To address this gap, we propose a novel class-incremental grouping network (CIGN) that can learn category-wise semantic features to achieve continual audio-visual learning. Our CIGN leverages learnable audio-visual class tokens and audio-visual grouping to continually aggregate class-aware features. Additionally, it utilizes class tokens distillation and continual grouping to prevent forgetting parameters learned from previous tasks, thereby improving the model’s ability to capture discriminative audiovisual categories. We conduct extensive experiments on VGG-Sound-Instruments, VGGSound-100, and VGG-Sound Sources benchmarks. Our experimental results demonstrate that the CIGN achieves state-of-the-art audio-visual class-incremental learning performance. Code is available at https://github.com/stoneMo/CIGN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Weakly-Supervised Audio-Visual SegmentationShentong Mo, Bhiksha RajNeurIPS 2023 · 被引用 26 次
- Continual Audio-Visual Sound SeparationWeiguo Pian, Yiyang Nan, Shijian Deng, Shentong Mo 等NeurIPS 2024 · 被引用 11 次
- Aligning Audio-Visual Joint Representations with an Agentic WorkflowShentong Mo, Yibing SongNeurIPS 2024 · 被引用 7 次
- PreFM: Online Audio-Visual Event Parsing via Predictive Future ModelingXiao Yu, Yan Fang, Yao Zhao, Yunchao WeiNeurIPS 2025 · 被引用 4 次
- Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and RegularizationYawen Cui, Li Liu, Zitong Yu, Guanjie Huang 等AAAI 2025 · 被引用 2 次
它引用的顶会 Paper30
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- Co2L: Contrastive Continual LearningHyuntak Cha, Jaeho Lee, Jinwoo ShinICCV 2021 · 被引用 391 次
相关 Paper
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
- Audio-Visual Grouping Network for Sound Localization from MixturesShentong Mo, Yapeng TianCVPR 2023
- AttriCLIP: A Non-Incremental Learner for Incremental Knowledge LearningRunqi Wang, Xiaoyue Duan, Guoliang Kang, Jianzhuang Liu 等CVPR 2023
- Knowledge Graph Enhanced Generative Multi-modal Models for Class-Incremental LearningXusheng Cao, Haori Lu, Linlan Huang, Fei Yang 等NeurIPS 2025 · 被引用 3 次
- AVQACL: A Novel Benchmark for Audio-Visual Question Answering Continual LearningKaixuan Wu, Xinde Li, Xinling Li, Chuanfei Hu 等CVPR 2025
