FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech Recognition
Yichong Leng, Xu Tan, Linchen Zhu, Jin Xu, Renqian Luo, Linquan Liu, Tao Qin, Xiangyang Li, Edward Lin, Tie-Yan Liu
Abstract
Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER) than original ASR outputs. Previous works usually use a sequence-to-sequence model to correct an ASR output sentence autoregressively, which causes large latency and cannot be deployed in online ASR services. A straightforward solution to reduce latency, inspired by non-autoregressive (NAR) neural machine translation, is to use an NAR sequence generation model for ASR error correction, which, however, comes at the cost of significantly increased ASR error rate. In this paper, observing distinctive error patterns and correction operations (i.e., insertion, deletion, and substitution) in ASR, we propose FastCorrect, a novel NAR error correction model based on edit alignment. In training, FastCorrect aligns each source token from an ASR output sentence to the target tokens from the corresponding ground-truth sentence based on the edit distance between the source and target sentences, and extracts the number of target tokens corresponding to each source token during edition/correction, which is then used to train a length predictor and to adjust the source tokens to match the length of the target sentence for parallel generation. In inference, the token number predicted by the length predictor is used to adjust the source tokens for target sequence generation. Experiments on the public AISHELL-1 dataset and an internal industrial-scale ASR dataset show the effectiveness of FastCorrect for ASR error correction: 1) it speeds up the inference by 6-9 times and maintains the accuracy (8-14% WER reduction) compared with the autoregressive correction model; and 2) it outperforms the popular NAR models adopted in neural machine translation and text edition by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19850a7f-927d-4f13-a965-1b264bbcaaf8Cited by top-tier papers10
- SoftCorrect: Error Correction with Soft Detection for Automatic Speech RecognitionYichong Leng, Xu Tan, Wenjie Liu, Kaitao Song et al.AAAI 2023 · 22 citations
- GenTranslate: Large Language Models are Generative Multilingual Speech and Machine TranslatorsYuchen Hu, Chen Chen, Chao-Han Huck Yang, Ruizhe Li et al.ACL 2024 · 15 citations
- Transcormer: Transformer for Sentence Scoring with Sliding Language ModelingKaitao Song, Yichong Leng, Xu Tan, Yicheng Zou et al.NeurIPS 2022 · 12 citations
- Mask the Correct Tokens: An Embarrassingly Simple Approach for Error CorrectionKai Shen, Yichong Leng, Xu Tan, Siliang Tang et al.EMNLP 2022 · 8 citations
- MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-FormulaSieun Hyeon, Kyudan Jung, Jaehee Won, Nam-Joon Kim et al.AAAI 2025 · 7 citations
Builds on5
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta PosteriorRaphael Shu, Jason Lee, Hideki Nakayama, Kyunghyun ChoAAAI 2020 · 125 citations
- Aligned Cross Entropy for Non-Autoregressive Machine TranslationMarjan Ghazvininejad, Vladimir Karpukhin, Luke Zettlemoyer, Omer LevyICML 2020 · 121 citations
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine TranslationChenze Shao, Jinchao Zhang, Yang Feng, Fandong Meng et al.AAAI 2020 · 95 citations
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine TranslationJunliang Guo, Xu Tan, Linli Xu, Tao Qin et al.AAAI 2020 · 91 citations
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao et al.ACL 2020 · 58 citations
Related papers
- NAT4AT: Using Non-Autoregressive Translation Makes Autoregressive Translation Faster and BetterHuanran Zheng, Wei Zhu, Xiaoling WangWWW 2024 · 13 citations
- An EM Approach to Non-autoregressive Conditional Sequence GenerationZhiqing Sun, Yiming YangICML 2020 · 43 citations
- SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative DecodingLinye Wei, Shuzhang Zhong, Songqiang Xu, Runsheng Wang et al.DAC 2025 · 7 citations
- NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic LanguagesLakshya Tomar, Vinayak Abrol, Puneet AgarwalAAAI 2026
- Improving Non-Autoregressive Translation Models Without DistillationXiao Shi Huang, Felipe Pérez, Maksims VolkovsICLR 2022 · 60 citations
