Unsupervised Speech Recognition
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, Michael Auli
Abstract
Despite rapid progress in the recent past, current speech recognition systems still require labeled training data which limits this technology to a small fraction of the languages spoken around the globe. This paper describes wav2vec-U, short for wav2vec Unsupervised, a method to train speech recognition models without any labeled data. We leverage self-supervised speech representations to segment unlabeled audio and learn a mapping from these representations to phonemes via adversarial training. The right representations are key to the success of our method. Compared to the best previous unsupervised work, wav2vec-U reduces the phoneme error rate on the TIMIT benchmark from 26.1 to 11.3. On the larger English Librispeech benchmark, wav2vec-U achieves a word error rate of 5.9 on test-other, rivaling some of the best published systems trained on 960 hours of labeled data from only two years ago. We also experiment on nine other languages, including low-resource languages such as Kyrgyz, Swahili and Tatar.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a2d85b51-a301-4e20-939c-f8d7e18ca2a8Cited by top-tier papers21
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and LanguageAlexei Baevski, Arun Babu, Wei-Ning Hsu, Michael AuliICML 2023 · 137 citations
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionCheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang et al.NeurIPS 2021 · 91 citations
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song et al.NeurIPS 2022 · 81 citations
Builds on3
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded SpeechDavid Harwath, Wei-Ning Hsu, James R. GlassICLR 2020 · 88 citations
Related papers
- Self-supervised learning with random-projection quantizer for speech recognitionChung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu et al.ICML 2022 · 245 citations
- Simple and Effective Unsupervised Speech TranslationChanghan Wang, Hirofumi Inaguma, Peng-Jen Chen, Ilia Kulikov et al.ACL 2023 · 9 citations
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani et al.ICML 2021 · 140 citations
- From perception to production: how acoustic invariance facilitates articulatory learning in a self-supervised vocal imitation modelMarvin Lavechin, Thomas HueberEMNLP 2025
- Toward a realistic model of speech processing in the brain with self-supervised learningJuliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec et al.NeurIPS 2022 · 164 citations
