Word-Level Speech Recognition With a Letter to Word Encoder
Ronan Collobert, Awni Y. Hannun, Gabriel Synnaeve
摘要
We propose a direct-to-word sequence model which uses a word network to learn word embeddings from letters. The word network can be integrated seamlessly with arbitrary sequence models including Connectionist Temporal Classification and encoder-decoder models with attention. We show our direct-to-word model can achieve word error rate gains over sub-word level models for speech recognition. We also show that our direct-to-word approach retains the ability to predict words not seen at training time without any retraining. Finally, we demonstrate that a word-level model can use a larger stride than a sub-word level model while maintaining accuracy. This makes the model more efficient both for training and inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- EM-Network: Oracle Guided Self-distillation for Sequence LearningJi Won Yoon, Sunghwan Ahn, Hyeonseung Lee, Minchan Kim 等ICML 2023 · 被引用 3 次
- Bridging the Gap between Pre-Training and Fine-Tuning for End-to-End Speech TranslationChengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang 等AAAI 2020 · 被引用 90 次
- W-CTC: a Connectionist Temporal Classification Loss with Wild CardsXingyu Cai, Jiahong Yuan, Yuchen Bian, Guangxu Xun 等ICLR 2022 · 被引用 16 次
- Align With Purpose: Optimize Desired Properties in CTC Models with a General Plug-and-Play FrameworkEliya Segev, Maya Alroy, Ronen Katsir, Noam Wies 等ICLR 2024 · 被引用 2 次
- T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask LearningNabarun Goswami, Hanqin Wang, Tatsuya HaradaICLR 2025
