Measuring the Effect of Transcription Noise on Downstream Language Understanding Tasks
Ori Shapira, Shlomo E. Chazan, Amir David Nissan Cohen
Abstract
With the increasing prevalence of recorded human speech, spoken language understanding (SLU) is essential for its efficient processing. In order to process the speech, it is commonly transcribed using automatic speech recognition technology. This speech-to-text transition introduces errors into the transcripts, which subsequently propagate to downstream NLP tasks, such as dialogue summarization. While it is known that transcript noise affects downstream tasks, a systematic approach to analyzing its effects across different noise severities and types has not been addressed. We propose a configurable framework for assessing task models in diverse noisy settings, and for examining the impact of transcript-cleaning techniques. The framework facilitates the investigation of task model behavior, which can in turn support the development of effective SLU solutions. We exemplify the utility of our framework on three SLU tasks and four task models, offering insights regarding the effect of transcript noise on tasks in general and models in particular. For instance, we find that task models can tolerate a certain level of noise, and are affected differently by the types of errors in the transcript. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0334bfda-6969-4c2e-81bc-2e739df8e1f8Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- QAConv: Question Answering on Informative ConversationsChien-Sheng Wu, Andrea Madotto, Wenhao Liu, Pascale Fung et al.ACL 2022 · 34 citations
- SLUE Phase-2: A Benchmark Suite of Diverse Spoken Language Understanding TasksSuwon Shon, Siddhant Arora, Chyi-Jiunn Lin, Ankita Pasad et al.ACL 2023 · 21 citations
- Using Phoneme Representations to Build Predictive Models Robust to ASR ErrorsAnjie Fang, Simone Filice, Nut Limsopatham, Oleg RokhlenkoSIGIR 2020 · 16 citations
- Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition ErrorsMarek Kubis, Pawel Skórzewski, Marcin Sowanski, Tomasz ZietkiewiczEMNLP 2023 · 7 citations
Related papers
- Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingYeonJoon Jung, Jaeseong Lee, Seungtaek Choi, Dohyeon Lee et al.EMNLP 2024
- DialogLM: Pre-trained Model for Long Dialogue Understanding and SummarizationMing Zhong, Yang Liu, Yichong Xu, Chenguang Zhu et al.AAAI 2022 · 150 citations
- Robustness Testing of Language Understanding in Task-Oriented DialogJiexi Liu, Ryuichi Takanobu, Jiaxin Wen, Dazhen Wan et al.ACL 2021
- MEDSAGE: Enhancing Robustness of Medical Dialogue Summarization to ASR Errors with LLM-generated Synthetic DialoguesKuluhan Binici, Abhinav Ramesh Kashyap, Viktor Schlegel, Andy T. Liu et al.AAAI 2025 · 10 citations
- Why Aren't We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech TranscriptsPiotr Szymanski, Lukasz Augustyniak, Mikolaj Morzy, Adrian Szymczak et al.ACL 2023 · 7 citations
