UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled Data
Chengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani, Shujie Liu, Furu Wei, Michael Zeng, Xuedong Huang
Abstract
In this paper, we propose a unified pre-training approach called UniSpeech to learn speech representations with both unlabeled and labeled data, in which supervised phonetic CTC learning and phonetically-aware contrastive self-supervised learning are conducted in a multi-task learning manner. The resultant representations can capture information more correlated with phonetic structures and improve the generalization across languages and domains. We evaluate the effectiveness of UniSpeech for cross-lingual representation learning on public CommonVoice corpus. The results show that UniSpeech outperforms self-supervised pretraining and supervised transfer learning for speech recognition by a maximum of 13.4% and 17.8% relative phone error rate reductions respectively (averaged over all testing languages). The transferability of UniSpeech is also demonstrated on a domain-shift speech recognition task, i.e., a relative word error rate reduction of 6% against the previous approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c280738-516a-4401-981a-46c033072ebbCited by top-tier papers8
- Squeezeformer: An Efficient Transformer for Automatic Speech RecognitionSehoon Kim, Amir Gholami, Albert E. Shaw, Nicholas Lee et al.NeurIPS 2022 · 152 citations
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech RecognitionCheng-I Jeff Lai, Yang Zhang, Alexander H. Liu, Shiyu Chang et al.NeurIPS 2021 · 91 citations
- Achieving Cross Modal Generalization with Multimodal Unified RepresentationYan Xia, Hai Huang, Jieming Zhu, Zhou ZhaoNeurIPS 2023 · 84 citations
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer et al.ICLR 2023 · 54 citations
- Finding Order in Chaos: A Novel Data Augmentation Method for Time Series in Contrastive LearningBerken Utku Demirel, Christian HolzNeurIPS 2023 · 48 citations
Builds on1
Related papers
- Unified Speech-Text Pre-training for Speech Translation and RecognitionYun Tang, Hongyu Gong, Ning Dong, Changhan Wang et al.ACL 2022 · 104 citations
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang et al.ACL 2023 · 13 citations
- SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language ProcessingJunyi Ao, Rui Wang, Long Zhou, Chengyi Wang et al.ACL 2022
- GCC: Graph Contrastive Coding for Graph Neural Network Pre-TrainingJiezhong Qiu, Qibin Chen, Yuxiao Dong, Jing Zhang et al.KDD 2020 · 755 citations
- AudioMosaic: Contrastive Masked Audio Representation LearningHanxun Huang, Qizhou Wang, Xingjun Ma, Cihang Xie et al.ICML 2026 · 2 citations
