PACE: Pretrained Audio Continual Learning
Chang Li, Kanglei Zhou, Liyuan Wang
摘要
Audio is a fundamental modality for analyzing speech, music, and environmental sounds. While pretrained audio models have significantly advanced audio understanding, they remain fragile in real-world scenarios where data distributions evolve over time. In this work, we present the first systematic benchmark for audio continual learning (CL) with pretrained models (PTMs) and provide a comprehensive analysis of its unique challenges. Unlike in the vision domain where parameter-efficient fine-tuning (PEFT) has proven effective for CL, directly applying such strategies to audio leads to poor performance. This is due to a fundamental property of audio backbones: they emphasize low-level spectral details rather than structured semantics, resulting in severe upstream–downstream misalignment. Through extensive empirical analysis, we identify a promising technical route based on analytic classifiers with first-session adaptation (FSA), but also uncover two major limitations: representation saturation in coarse-grained scenarios and representation shifts in fine-grained scenarios. To address these challenges, we propose PACE, an innovative method that improves FSA via a regularized analytic classifier and introduces multi-session adaptation through adaptive subspace-orthogonal PEFT for better semantic alignment. Additionally, we design spectrogram-based boundary-aware perturbations to mitigate representation overlap and improve stability. Experiments across six diverse audio CL benchmarks demonstrate that PACE substantially outperforms state-of-the-art baselines, representing a significant step toward robust and scalable audio CL with PTMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- S-Prompts Learning with Pre-trained Transformers: An Occam's Razor for Domain Incremental LearningYabin Wang, Zhiwu Huang, Xiaopeng HongNeurIPS 2022 · 被引用 397 次
- SSAST: Self-Supervised Audio Spectrogram TransformerYuan Gong, Cheng-I Lai, Yu-An Chung, James R. GlassAAAI 2022 · 被引用 397 次
相关 Paper
- TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio ClassificationYunlong Gao, Wenxin Liang, Guanglu Wang, Senqi Guan 等CVPR 2026
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding 等AAAI 2026
- Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual LearningLingfeng He, De Cheng, Di Xu, Huaijie Wang 等AAAI 2026 · 被引用 1 次
- Revisiting Audio-language Pretraining for Learning General-purpose Audio RepresentationWei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao 等ACL 2026 · 被引用 2 次
- Do You Remember? Overcoming Catastrophic Forgetting for Fake Audio DetectionXiaohui Zhang, Jiangyan Yi, Jianhua Tao, Chenglong Wang 等ICML 2023 · 被引用 33 次
