Global Normalization for Streaming Speech Recognition in a Modular Framework
Ehsan Variani, Ke Wu, Michael D. Riley, David Rybach, Matt Shannon, Cyril Allauzen
Abstract
We introduce the Globally Normalized Autoregressive Transducer (GNAT) for addressing the label bias problem in streaming speech recognition. Our solution admits a tractable exact computation of the denominator for the sequence-level normalization. Through theoretical and empirical results, we demonstrate that by switching to a globally normalized model, the word error rate gap between streaming and non-streaming speech-recognition models can be greatly reduced (by more than 50% on the Librispeech dataset). This model is developed in a modular framework which encompasses all the common neural speech recognition models. The modularity of this framework enables controlled comparison of modelling choices and creation of new models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38fca7ca-0a0f-44b2-b307-f4c02a6b28e5Cited by top-tier papers2
- Preference Optimization for Molecule Synthesis with Conditional Residual Energy-based ModelsSongtao Liu, Hanjun Dai, Yue Zhao, Peng LiuICML 2024 · 8 citations
- AdaStreamLite: Environment-adaptive Streaming Speech Recognition on Mobile DevicesYuheng Wei, Jie Xiong, Hui Liu, Yingtao Yu et al.UbiComp 2024 · 5 citations
Related papers
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar et al.ICML 2023 · 12 citations
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context ModelingJiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu et al.ICLR 2021 · 80 citations
- Imputer: Sequence Modelling via Imputation and Dynamic ProgrammingWilliam Chan, Chitwan Saharia, Geoffrey E. Hinton, Mohammad Norouzi et al.ICML 2020 · 127 citations
- Speech-T: Transducer for Text to Speech and BeyondJiawei Chen, Xu Tan, Yichong Leng, Jin Xu et al.NeurIPS 2021 · 23 citations
- Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation ModelsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Chengwei Qin et al.NeurIPS 2024 · 14 citations
