Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
Jiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu, Bo Li, Tara N. Sainath, Yonghui Wu, Ruoming Pang
摘要
Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full speech utterance before emitting completed hypotheses. In this work, we propose a unified framework, Dual-mode ASR, to train a single end-to-end ASR model with shared weights for both streaming and full-context speech recognition. We show that the latency and accuracy of streaming ASR significantly benefit from weight sharing and joint training of full-context ASR, especially with inplace knowledge distillation during the training. The Dual-mode ASR framework can be applied to recent state-of-the-art convolution-based and transformer-based ASR networks. We present extensive experiments with two state-of-the-art ASR networks, ContextNet and Conformer, on two datasets, a widely used public dataset LibriSpeech and a large-scale dataset MultiDomain. Experiments and ablation studies demonstrate that Dual-mode ASR not only simplifies the workflow of training and deploying streaming and full-context ASR models, but also significantly improves both emission latency and recognition accuracy of streaming ASR. With Dual-mode ASR, we achieve new state-of-the-art streaming ASR results on both LibriSpeech and MultiDomain in terms of accuracy and latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionCheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang 等NeurIPS 2021 · 被引用 91 次
- Unified Segment-to-Segment Framework for Simultaneous Sequence GenerationShaolei Zhang, Yang FengNeurIPS 2023 · 被引用 9 次
- Revisiting the Entropy Semiring for Neural Speech RecognitionOscar Chang, Dongseong Hwang, Olivier SiohanICLR 2023 · 被引用 1 次
- A Full-duplex Speech Dialogue Scheme Based On Large Language ModelPeng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan 等NeurIPS 2024
- An Efficient Self-Learning Framework For Interactive Spoken Dialog SystemsHitesh Tulsiani, David M. Chan, Shalini Ghosh, Garima Lalwani 等ICML 2024
它引用的顶会 Paper3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
相关 Paper
- Heuristic-free Knowledge Distillation for Streaming ASR via Multi-modal TrainingJi Won YoonAAAI 2025
- Radio2Text: Streaming Speech Recognition Using mmWave Radio SignalsRunning Zhao, Jiangtao Yu, Hang Zhao, Edith C. H. NgaiUbiComp 2023 · 被引用 22 次
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar 等ICML 2023 · 被引用 12 次
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu 等NeurIPS 2021 · 被引用 23 次
- VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelZuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao 等NeurIPS 2025 · 被引用 6 次
