Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognition
Yanshuo Wang, Yanghao Zhou, Yukang Lin, Haoxing Chen, Jin Zhang, Wentao Zhu, Jie Hong, Xuesong Li
摘要
End-to-end automatic speech recognition (ASR) based on deep learning has achieved impressive progress in recent years. However, the performance of ASR foundation model often degrades significantly on out-of-domain data due to real-world domain shifts. Test-Time Adaptation (TTA) methods aim to mitigate this issue by adapting models during inference without access to source data. Despite recent progress, existing ASR TTA methods often struggle with instability under continual and long-term distribution shifts. To alleviate the risk of performance collapse due to error accumulation, we propose Dynamic Model-bank Single-Utterance Test-time Adaptation (DM-SUTA), a sustainable continual TTA framework based on adaptive ASR model ensembling. DMSUTA maintains a dynamic model bank, from which a subset of checkpoints is selected for each test sample based on confidence and uncertainty criteria. To preserve both model plasticity and long-term stability, DMSUTA actively manages the bank by filtering out potentially collapsed models. This design allows DMSUTA to continually adapt to evolving domain shifts in ASR test-time scenarios. Experiments on diverse, continuously shifting ASR TTA benchmarks show that DM-SUTA consistently outperforms existing continual TTA baselines, demonstrating superior robustness to domain shifts in ASR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen 等ICML 2022 · 被引用 579 次
- ViDA: Homeostatic Visual Domain Adapter for Continual Test Time AdaptationJiaming Liu, Senqiao Yang, Peidong Jia, Renrui Zhang 等ICLR 2024 · 被引用 71 次
相关 Paper
- Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy SpeechGuan-Ting Lin, Wei Huang, Hung-yi LeeEMNLP 2024 · 被引用 3 次
- Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic RewardsLinghan Fang, Tianxin Xie, Li LiuAAAI 2026 · 被引用 1 次
- ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring DomainsGuillaume Vray, Devavrat Tomar, Xufeng Gao, Jean-Philippe Thiran 等NeurIPS 2025 · 被引用 9 次
- Advancing Test-Time Adaptation in Wild Acoustic Test SettingsHongfu Liu, Hengguan Huang, Ye WangEMNLP 2024 · 被引用 2 次
- E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation ModelsJiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh 等NeurIPS 2025 · 被引用 10 次
