Dynamic Model-Bank Test-Time Adaptation for Automatic Speech Recognition
Yanshuo Wang, Yanghao Zhou, Yukang Lin, Haoxing Chen, Jin Zhang, Wentao Zhu, Jie Hong, Xuesong Li
Abstract
End-to-end automatic speech recognition (ASR) based on deep learning has achieved impressive progress in recent years. However, the performance of ASR foundation model often degrades significantly on out-of-domain data due to real-world domain shifts. Test-Time Adaptation (TTA) methods aim to mitigate this issue by adapting models during inference without access to source data. Despite recent progress, existing ASR TTA methods often struggle with instability under continual and long-term distribution shifts. To alleviate the risk of performance collapse due to error accumulation, we propose Dynamic Model-bank Single-Utterance Test-time Adaptation (DM-SUTA), a sustainable continual TTA framework based on adaptive ASR model ensembling. DMSUTA maintains a dynamic model bank, from which a subset of checkpoints is selected for each test sample based on confidence and uncertainty criteria. To preserve both model plasticity and long-term stability, DMSUTA actively manages the bank by filtering out potentially collapsed models. This design allows DMSUTA to continually adapt to evolving domain shifts in ASR test-time scenarios. Experiments on diverse, continuously shifting ASR TTA benchmarks show that DM-SUTA consistently outperforms existing continual TTA baselines, demonstrating superior robustness to domain shifts in ASR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45415069-d281-416b-801c-21304a4dfd8fBuilds on9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Efficient Test-Time Model Adaptation without ForgettingShuaicheng Niu, Jiaxiang Wu, Yifan Zhang, Yaofo Chen et al.ICML 2022 · 579 citations
- ViDA: Homeostatic Visual Domain Adapter for Continual Test Time AdaptationJiaming Liu, Senqiao Yang, Peidong Jia, Renrui Zhang et al.ICLR 2024 · 71 citations
Related papers
- Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy SpeechGuan-Ting Lin, Wei Huang, Hung-yi LeeEMNLP 2024 · 3 citations
- Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic RewardsLinghan Fang, Tianxin Xie, Li LiuAAAI 2026 · 1 citation
- ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring DomainsGuillaume Vray, Devavrat Tomar, Xufeng Gao, Jean-Philippe Thiran et al.NeurIPS 2025 · 9 citations
- Advancing Test-Time Adaptation in Wild Acoustic Test SettingsHongfu Liu, Hengguan Huang, Ye WangEMNLP 2024 · 2 citations
- E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation ModelsJiaheng Dong, Hong Jia, Soumyajit Chatterjee, Abhirup Ghosh et al.NeurIPS 2025 · 10 citations
