CopyNE: Better Contextual ASR by Copying Named Entities
Shilin Zhou, Zhenghua Li, Yu Hong, Min Zhang, Zhefeng Wang, Baoxing Huai
Abstract
End-to-end automatic speech recognition (ASR) systems have made significant progress in general scenarios. However, it remains challenging to transcribe contextual named entities (NEs) in the contextual ASR scenario. Previous approaches have attempted to address this by utilizing the NE dictionary. These approaches treat entities as individual tokens and generate them token-by-token, which may result in incomplete transcriptions of entities. In this paper, we treat entities as indivisible wholes and introduce the idea of copying into ASR. We design a systematic mechanism called CopyNE, which can copy entities from the NE dictionary. By copying all tokens of an entity at once, we can reduce errors during entity transcription, ensuring the completeness of the entity. Experiments demonstrate that CopyNE consistently improves the accuracy of transcribing entities compared to previous approaches. Even when based on the strong Whisper, CopyNE still achieves notable improvements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3042d2c-8f65-4d5e-bd9d-7c4b3b1f71adBuilds on5
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- SpeechUT: Bridging Speech and Text with Hidden-Unit for Encoder-Decoder Based Speech-Text Pre-trainingZiqiang Zhang, Long Zhou, Junyi Ao, Shujie Liu et al.EMNLP 2022 · 38 citations
- Towards relation extraction from speechTongtong Wu, Guitao Wang, Jinming Zhao, Zhaoran Liu et al.EMNLP 2022 · 6 citations
- Copy is All You NeedTian Lan, Deng Cai, Yan Wang, Heyan Huang et al.ICLR 2023
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang et al.ACL 2022
Related papers
- Generative Annotation for ASR Named Entity CorrectionYuanchang Luo, Daimeng Wei, Shaojun Li, Hengchao Shang et al.EMNLP 2025
- Why Aren't We NER Yet? Artifacts of ASR Errors in Named Entity Recognition in Spontaneous Speech TranscriptsPiotr Szymanski, Lukasz Augustyniak, Mikolaj Morzy, Adrian Szymczak et al.ACL 2023 · 7 citations
- Speech-enriched Memory for Inference-time Adaptation of ASR Models to Word DictionariesAshish R. Mittal, Sunita Sarawagi, Preethi Jyothi, George Saon et al.EMNLP 2023 · 2 citations
- WhisperDiari: A Whisper-Based Speaker Diarization Framework in Token Space Leveraging Semantic and Speaker Information for Better Text AdaptabilityYongkang Yin, Yuexian ZouAAAI 2026
- LUKE: Deep Contextualized Entity Representations with Entity-aware Self-attentionIkuya Yamada, Akari Asai, Hiroyuki Shindo, Hideaki Takeda et al.EMNLP 2020 · 562 citations
