Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic Segmentation
Jingqiao Xiu, Mengze Li, Zongxin Yang, Wei Ji, Yifang Yin, Roger Zimmermann
摘要
Audio-Visual Semantic Segmentation (AVSS) has gained significant attention in the multi-modal domain, aiming to segment video objects that produce specific sounds in the corresponding audio. Despite notable progress, existing methods still struggle to handle new classes not included in the original training set. To this end, we introduce Few-Shot Incremental Learning (FSIL) to the AVSS task, which seeks to seamlessly integrate new classes with limited incremental samples while preserving the knowledge of old classes. Two challenges arise in this new setting: (1) To reduce labeling costs, old classes within the incremental samples are treated as background, similar to silent objects. Training the model directly with background annotations may worsen the loss of distinctive knowledge about old classes, such as their outlines and sounds. (2) Most existing models adopt early cross-modal fusion with a single-tower design, incorporating more characteristics into class representations, which impedes knowledge transfer between classes based on similarity. To address these issues, we propose a Few-shot Incremental learning framework via class-centric foregrouNd aggreGation and dual-tower knowlEdge tRansfer (FINGER) for the AVSS task, which comprises two targeted modules: (1) The class-centric foreground aggregation gathers class-specific features for each foreground class while disregarding background features. The background class is excluded during training and inferred from the foreground predictions. (2) The dual-tower knowledge transfer postpones cross-modal fusion to separately conduct knowledge transfer for each modality. Extensive experiments validate the effectiveness of the FINGER model, significantly surpassing state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object CategoriesYicong Li, Yiyang Chen, Zhenyuan Ma, Junbin Xiao 等ICCV 2025 · 被引用 3 次
- Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesJingqiao Xiu, Yicong Li, Na Zhao, Han Fang 等ICCV 2025 · 被引用 2 次
- MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and EditingJingqiao Xiu, Can Wang, Dong XuCVPR 2026
它引用的顶会 Paper32
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Associating Objects with Transformers for Video Object SegmentationZongxin Yang, Yunchao Wei, Yi YangNeurIPS 2021 · 被引用 398 次
- Dual Attention Matching for Audio-Visual Event LocalizationYu Wu, Linchao Zhu, Yan Yan, Yi YangICCV 2019 · 被引用 233 次
- Few-Shot Class-Incremental Learning via Relation Knowledge DistillationSonglin Dong, Xiaopeng Hong, Xiaoyu Tao, Xinyuan Chang 等AAAI 2021 · 被引用 215 次
- Constrained Few-shot Class-incremental LearningMichael Hersche, Geethan Karunaratne, Giovanni Cherubini, Luca Benini 等CVPR 2022 · 被引用 152 次
相关 Paper
- Few-Shot Audio-Visual Class-Incremental Learning with Temporal Prompting and RegularizationYawen Cui, Li Liu, Zitong Yu, Guanjie Huang 等AAAI 2025 · 被引用 2 次
- Incremental Few Shot Semantic Segmentation via Class-agnostic Mask Proposal and Language-driven ClassifierLeo Shan, Wenzhang Zhou, Grace ZhaoACM MM 2023 · 被引用 20 次
- Prototypical Kernel Learning and Open-set Foreground Perception for Generalized Few-shot Semantic SegmentationKai Huang, Feigege Wang, Ye Xi, Yutao GaoICCV 2023 · 被引用 16 次
- Audio-Visual Class-Incremental LearningWeiguo Pian, Shentong Mo, Yunhui Guo, Yapeng TianICCV 2023 · 被引用 44 次
- Integrative Few-Shot Learning for Classification and SegmentationDahyun Kang, Minsu ChoCVPR 2022 · 被引用 76 次
