Why Aren't We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech Transcripts
Piotr Szymanski, Lukasz Augustyniak, Mikolaj Morzy, Adrian Szymczak, Krzysztof Surdyk, Piotr Zelasko
摘要
Transcripts of spontaneous human speech present a significant obstacle for traditional NER models. The lack of grammatical structure of spoken utterances and word errors introduced by the ASR make downstream NLP tasks challenging. In this paper, we examine in detail the complex relationship between ASR and NER errors which limit the ability of NER models to recover entity mentions from spontaneous speech transcripts. Using publicly available benchmark datasets (SWNE, Earnings-21, OntoNotes), we present the full taxonomy of ASR-NER errors and measure their true impact on entity recognition. We find that NER models fail to recognize entity spans even if no word errors are introduced by the ASR. We also show why the F 1 score is inadequate to evaluate NER models on conversational transcripts 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Generative Annotation for ASR Named Entity CorrectionYuanchang Luo, Daimeng Wei, Shaojun Li, Hengchao Shang 等EMNLP 2025
- CopyNE: Better Contextual ASR by Copying Named EntitiesShilin Zhou, Zhenghua Li, Yu Hong, Min Zhang 等ACL 2024
- NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity RecognitionElena Merdjanovska, Ansar Aynetdinov, Alan AkbikEMNLP 2024 · 被引用 5 次
- Recording for Eyes, Not Echoing to Ears: Contextualized Spoken-to-Written Conversion of ASR TranscriptsJiaqing Liu, Chong Deng, Qinglin Zhang, Shilin Zhou 等AAAI 2025 · 被引用 1 次
- Using Phoneme Representations to Build Predictive Models Robust to ASR ErrorsAnjie Fang, Simone Filice, Nut Limsopatham, Oleg RokhlenkoSIGIR 2020 · 被引用 16 次
