SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition
Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Edward Lin, Tie-Yan Liu
Abstract
Error correction in automatic speech recognition (ASR) aims to correct those incorrect words in sentences generated by ASR models. Since recent ASR models usually have low word error rate (WER), to avoid affecting originally correct tokens, error correction models should only modify incorrect words, and therefore detecting incorrect words is important for error correction. Previous works on error correction either implicitly detect error words through target-source attention or CTC (connectionist temporal classification) loss, or explicitly locate specific deletion/substitution/insertion errors. However, implicit error detection does not provide clear signal about which tokens are incorrect and explicit error detection suffers from low detection accuracy. In this paper, we propose SoftCorrect with a soft error detection mechanism to avoid the limitations of both explicit and implicit error detection. Specifically, we first detect whether a token is correct or not through a probability produced by a dedicatedly designed language model, and then design a constrained CTC loss that only duplicates the detected incorrect tokens to let the decoder focus on the correction of error tokens. Compared with implicit error detection with CTC loss, SoftCorrect provides explicit signal about which words are incorrect and thus does not need to duplicate every token but only incorrect tokens; compared with explicit error detection, Soft-Correct does not detect specific deletion/substitution/insertion errors but just leaves it to CTC loss. Experiments on AISHELL-1 and Aidatatang datasets show that SoftCorrect achieves 26.1% and 9.4% CER reduction respectively, outperforming previous works by a large margin, while still enjoying fast speed of parallel generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89d15e7c-d0b2-40ef-9598-b4f27144524bCited by top-tier papers5
- It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech RecognitionChen Chen, Ruizhe Li, Yuchen Hu, Sabato Marco Siniscalchi et al.ICLR 2024 · 37 citations
- MathSpeech: Leveraging Small LMs for Accurate Conversion in Mathematical Speech-to-FormulaSieun Hyeon, Kyudan Jung, Jaehee Won, Nam-Joon Kim et al.AAAI 2025 · 7 citations
- WavePurifier: Purifying Audio Adversarial Examples via Hierarchical Diffusion ModelsHanqing Guo, Guangjing Wang, Bocheng Chen, Yuanda Wang et al.MobiCom 2024 · 3 citations
- A Simple yet Effective Training-free Prompt-free Approach to Chinese Spelling Correction Based on Large Language ModelsHouquan Zhou, Zhenghua Li, Bo Zhang, Chen Li et al.EMNLP 2024 · 2 citations
- Optimized Tokenization for Transcribed Error CorrectionTomer Wullach, Shlomo E. ChazanEMNLP 2023 · 1 citation
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- FastCorrect: Fast Error Correction with Edit Alignment for Automatic Speech RecognitionYichong Leng, Xu Tan, Linchen Zhu, Jin Xu et al.NeurIPS 2021 · 84 citations
- Mask the Correct Tokens: An Embarrassingly Simple Approach for Error CorrectionKai Shen, Yichong Leng, Xu Tan, Siliang Tang et al.EMNLP 2022 · 8 citations
- Non-Autoregressive Machine Translation with Latent AlignmentsChitwan Saharia, William Chan, Saurabh Saxena, Mohammad NorouziEMNLP 2020 · 6 citations
Related papers
- Spelling Error Correction with Soft-Masked BERTShaohua Zhang, Haoran Huang, Jicong Liu, Hang LiACL 2020 · 204 citations
- Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play FrameworkEliya Segev, Maya Alroy, Ronen Katsir, Noam Wies et al.ICLR 2024 · 2 citations
- UMRSpell: Unifying the Detection and Correction Parts of Pre-trained Models towards Chinese Missing, Redundant, and Spelling CorrectionZheyu He, Yujin Zhu, Linlin Wang, Liang XuACL 2023 · 8 citations
- W-CTC: a Connectionist Temporal Classification Loss with Wild CardsXingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun et al.ICLR 2022 · 16 citations
- CR-CTC: Consistency regularization on CTC for improved speech recognitionZengwei Yao, Wei Kang, Xiaoyu Yang, Fangjun Kuang et al.ICLR 2025
