ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David D. Cox, Mark Hasegawa-Johnson, Shiyu Chang
Abstract
Self-supervised learning (SSL) in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of the downstream tasks of SSL learning in speech largely focus on the content information in speech, the most desirable speech representations should be able to disentangle unwanted variations, such as speaker variations, from the content. However, disentangling speakers is very challenging, because removing the speaker information could easily result in a loss of content as well, and the damage of the latter usually far outweighs the benefit of the former. In this paper, we propose a new SSL method that can achieve speaker disentanglement without severe loss of content. Our approach is adapted from the HuBERT framework, and incorporates disentangling mechanisms to regularize both the teachers (masked prediction labels) and the students (learned representations). We evaluate the benefit of speaker disentanglement on a set of content-related downstream tasks, and observe a consistent and notable performance advantage of our speaker-disentangled representations. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ade335e6-bc5e-423e-acb3-d9f98e8d33cfCited by top-tier papers18
- Disentangling Voice and Content with Self-Supervision for Speaker RecognitionTianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou LiNeurIPS 2023 · 53 citations
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation LearningAlexander H. Liu, Heng-Jui Chang, Michael Auli, Wei-Ning Hsu et al.NeurIPS 2023 · 51 citations
- Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit PredictionJiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov et al.ICLR 2024 · 39 citations
- SSDM: Scalable Speech Dysfluency ModelingJiachen Lian, Xuanru Zhou, Zoe Ezzes, Jet Vonk et al.NeurIPS 2024 · 26 citations
- Self-Supervised Disentangled Representation Learning for Robust Target Speech ExtractionZhaoxi Mu, Xinyu Yang, Sining Sun, Qing YangAAAI 2024 · 13 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee et al.NeurIPS 2021 · 200 citations
Related papers
- Self-supervised Neural Factor Analysis for Disentangling Utterance-level Speech RepresentationsWeiwei Lin, Chenhang He, Man-Wai Mak, Youzhi TuICML 2023 · 6 citations
- SpeechTripleNet: End-to-End Disentangled Speech Representation Learning for Content, Timbre and ProsodyHui Lu, Xixin Wu, Zhiyong Wu, Helen MengACM MM 2023 · 5 citations
- Cross-modal Self-Supervised Learning for Lip Reading: When Contrastive Learning meets Adversarial TrainingChangchong Sheng, Matti Pietikäinen, Qi Tian, Li LiuACM MM 2021 · 11 citations
- Speech Self-Supervised Learning Using Diffusion Model Synthetic DataHeting Gao, Kaizhi Qian, Junrui Ni, Chuang Gan et al.ICML 2024 · 8 citations
- SelfVC: Voice Conversion With Iterative Refinement using Self TransformationsPaarth Neekhara, Shehzeen Samarah Hussain, Rafael Valle, Boris Ginsburg et al.ICML 2024 · 7 citations
