Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
Deepak Kumar, Baban Gain, Asif Ekbal
Abstract
Automatic Speech Recognition (ASR) transcripts often contain disfluencies, such as fillers, repetitions, and false starts, which reduce readability and hinder downstream applications like chatbots and voice assistants. If left unaddressed, such disfluencies can significantly degrade the reliability of downstream systems. Most existing approaches rely on classical models that focus on identifying disfluent tokens for removal. While this strategy is effective to some extent, it often disrupts grammatical structure and semantic coherence, leading to incomplete or unnatural sentences. Recent literature explored the use of large language models (LLMs); however, these efforts have primarily focused on disfluency detection or data augmentation, rather than performing comprehensive correction. We propose a multilingual correction pipeline where a sequence tagger first marks disfluent tokens, and these signals guide instruction fine-tuning of an LLM to rewrite transcripts into fluent text. To further improve reliability, we add a contrastive learning objective that penalizes the reproduction of disfluent tokens, encouraging the model to preserve grammar and meaning while removing disfluent artifacts. Our experiments across three Indian languages, namely Hindi, Bengali, and Marathi show consistent improvements over strong baselines, including multilingual sequence-to-sequence models. These results highlight that detection-only strategies are insufficient. Combining token-level cues with instruction tuning and contrastive learning provides a practical and scalable solution for multilingual disfluency correction in speechdriven NLP systems. We make the codes publicly available at https://github.com/ deepak-kumar-98/Mind-the-Pause .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- A Contrastive Framework for Neural Text GenerationYixuan Su, Tian Lan, Yan Wang, Dani Yogatama et al.NeurIPS 2022 · 349 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
Related papers
- LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form SpeechBingshen Mu, Xian Shi, Xiong Wang, Hexin Liu et al.ACL 2026 · 5 citations
- Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsJiaqing Liu, Chong Deng, Qinglin Zhang, Shilin Zhou et al.AAAI 2025 · 1 citation
- RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal AnalysisEnzhi Wang, Jiaming Zhou, Yuhang Jia, Aobo Kong et al.ACL 2026
- From English to Second Language Mastery: Enhancing LLMs with Cross-Lingual Continued Instruction TuningLinjuan Wu, Haoran Wei, Baosong Yang, Weiming LuACL 2025 · 3 citations
- Best of Both Worlds: Making High Accuracy Non-incremental Transformer-based Disfluency Detection IncrementalMorteza Rohanian, Julian HoughACL 2021
