HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
Hainan Xu, Travis M. Bartley, Vladimir Bataev, Boris Ginsburg
摘要
We present Hybrid-Autoregressive INference TrANsducers (HAINAN), a novel architecture for speech recognition that extends the Token-and-Duration Transducer (TDT) model. Trained with randomly masked predictor network outputs, HAINAN supports both autoregressive inference with all network components and non-autoregressive inference without the predictor. Additionally, we propose a novel semi-autoregressive inference method that first generates an initial hypothesis using non-autoregressive inference, followed by refinement steps where each token prediction is regenerated using parallelized autoregression on the initial hypothesis. Experiments on multiple datasets across different languages demonstrate that HAINAN achieves efficiency parity with CTC in non-autoregressive mode and with TDT in autoregressive mode. In terms of accuracy, autoregressive HAINAN achieves parity with TDT and RNN-T, while non-autoregressive HAINAN significantly outperforms CTC. Semi-autoregressive inference further enhances the model's accuracy with minimal computational overhead, and even outperforms TDT results in some cases. These results highlight HAINAN's flexibility in balancing accuracy and speed, positioning it as a strong candidate for real-world speech recognition applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang 等ICLR 2024 · 被引用 155 次
- Efficient Sequence Transduction by Jointly Predicting Tokens and DurationsHainan Xu, Fei Jia, Somshubra Majumdar, He Huang 等ICML 2023 · 被引用 62 次
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun 等ACL 2023 · 被引用 4 次
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu 等ACL 2021
相关 Paper
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng 等ICLR 2025
- Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersAdam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno MengibarNeurIPS 2024 · 被引用 4 次
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 被引用 42 次
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar 等ICML 2023 · 被引用 12 次
- Star Temporal Classification: Sequence Modeling with Partially Labeled DataVineel Pratap, Awni Hannun, Gabriel Synnaeve, Ronan CollobertNeurIPS 2022 · 被引用 7 次
