HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
Hainan Xu, Travis M. Bartley, Vladimir Bataev, Boris Ginsburg
Abstract
We present Hybrid-Autoregressive INference TrANsducers (HAINAN), a novel architecture for speech recognition that extends the Token-and-Duration Transducer (TDT) model. Trained with randomly masked predictor network outputs, HAINAN supports both autoregressive inference with all network components and non-autoregressive inference without the predictor. Additionally, we propose a novel semi-autoregressive inference method that first generates an initial hypothesis using non-autoregressive inference, followed by refinement steps where each token prediction is regenerated using parallelized autoregression on the initial hypothesis. Experiments on multiple datasets across different languages demonstrate that HAINAN achieves efficiency parity with CTC in non-autoregressive mode and with TDT in autoregressive mode. In terms of accuracy, autoregressive HAINAN achieves parity with TDT and RNN-T, while non-autoregressive HAINAN significantly outperforms CTC. Semi-autoregressive inference further enhances the model's accuracy with minimal computational overhead, and even outperforms TDT results in some cases. These results highlight HAINAN's flexibility in balancing accuracy and speed, positioning it as a strong candidate for real-world speech recognition applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Zipformer: A faster and better encoder for automatic speech recognitionZengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang et al.ICLR 2024 · 155 citations
- Efficient Sequence Transduction by Jointly Predicting Tokens and DurationsHainan Xu, Fei Jia, Somshubra Majumdar, He Huang et al.ICML 2023 · 62 citations
- CTC-based Non-autoregressive Speech TranslationChen Xu, Xiaoqian Liu, Xiaowen Liu, Qingxuan Sun et al.ACL 2023 · 4 citations
- VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and InterpretationChanghan Wang, Morgane Rivière, Ann Lee, Anne Wu et al.ACL 2021
Related papers
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng et al.ICLR 2025
- Aligner-Encoders: Self-Attention Transformers Can Be Self-TransducersAdam Stooke, Rohit Prabhavalkar, Khe Chai Sim, Pedro Moreno MengibarNeurIPS 2024 · 4 citations
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 42 citations
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar et al.ICML 2023 · 12 citations
- Star Temporal Classification: Sequence Modeling with Partially Labeled DataVineel Pratap, Awni Hannun, Gabriel Synnaeve, Ronan CollobertNeurIPS 2022 · 7 citations
