Using Phoneme Representations to Build Predictive Models Robust to ASR Errors
Anjie Fang, Simone Filice, Nut Limsopatham, Oleg Rokhlenko
摘要
Even though Automatic Speech Recognition (ASR) systems significantly improved over the last decade, they still introduce a lot of errors when they transcribe voice to text. One of the most common reasons for these errors is phonetic confusion between similar-sounding expressions. As a result, ASR transcriptions often contain "quasi-oronyms", i.e., words or phrases that sound similar to the source ones, but that have completely different semantics (e.g., "win" instead of "when" or "accessible on defecting" instead of "accessible and affecting"). These errors significantly affect the performance of downstream Natural Language Understanding (NLU) models (e.g., intent classification, slot filling, etc.) and impair user experience. To make NLU models more robust to such errors, we propose novel phonetic-aware text representations. Specifically, we represent ASR transcriptions at the phoneme level, aiming to capture pronunciation similarities, which are typically neglected in word-level representations (e.g., word embeddings). To train and evaluate our phoneme representations, we generate noisy ASR transcriptions of four existing datasets - Stanford Sentiment Treebank, SQuAD, TREC Question Classification and Subjectivity Analysis - and show that common neural network architectures exploiting the proposed phoneme representations can effectively handle noisy transcriptions and significantly outperform state-of-the-art baselines. Finally, we confirm these results by testing our models on real utterances spoken to the Alexa virtual assistant.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Measuring the Effect of Transcription Noise on Downstream Language Understanding TasksOri Shapira, Shlomo E. Chazan, Amir David Nissan CohenACL 2025 · 被引用 3 次
- Phun-Bench: Evaluating LLMs on Phonological Understanding in ChineseXing Yue, Yongliang Shen, Weiming LuACL 2026
相关 Paper
- Introducing Semantics into Speech EncodersDerek Xu, Shuyan Dong, Changhan Wang, Suyoun Kim 等ACL 2023 · 被引用 2 次
- Life after Speech Recognition: Fuzzing Semantic Misinterpretation for Voice Assistant ApplicationsYangyong Zhang, Lei Xu, Abner Mendoza, Guangliang Yang 等NDSS 2019 · 被引用 60 次
- Why Aren't We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech TranscriptsPiotr Szymanski, Lukasz Augustyniak, Mikolaj Morzy, Adrian Szymczak 等ACL 2023 · 被引用 7 次
- Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingYeonJoon Jung, Jaeseong Lee, Seungtaek Choi, Dohyeon Lee 等EMNLP 2024
- Exploring Transfer Learning For End-to-End Spoken Language UnderstandingSubendhu Rongali, Beiye Liu, Liwei Cai, Konstantine Arkoudas 等AAAI 2021 · 被引用 26 次
