Disentangling Voice and Content with Self-Supervision for Speaker Recognition
Tianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou Li
Abstract
For speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with the use of three Gaussian inference layers, each consisting of a learnable transition model that extracts distinct speech components. Notably, a strengthened transition model is specifically designed to model complex speech dynamics. We also propose a self-supervision method to dynamically disentangle content without the use of labels other than speaker identities. The efficacy of the proposed framework is validated via experiments conducted on the VoxCeleb and SITW datasets with 9.56% and 8.24% average reductions in EER and minDCF, respectively. Since neither additional model training nor data is specifically needed, it is easily applicable in practical use.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d29ac772-f550-4a04-a39e-ad9edd7aeebbCited by top-tier papers1
Ask how each one uses itBuilds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
Related papers
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
- Self-Supervised Disentangled Representation Learning for Robust Target Speech ExtractionZhaoxi Mu, Xinyu Yang, Sining Sun, Qing YangAAAI 2024 · 13 citations
- SpeechTripleNet: End-to-End Disentangled Speech Representation Learning for Content, Timbre and ProsodyHui Lu, Xixin Wu, Zhiyong Wu, Helen MengACM MM 2023 · 5 citations
- Bitrate-Controlled Diffusion for Disentangling Motion and Content in VideoXiao Li, Qi Chen, Xiulian Peng, Kai Yu et al.ICCV 2025 · 1 citation
- S3VAE: Self-Supervised Sequential VAE for Representation Disentanglement and Data GenerationYizhe Zhu, Martin Renqiang Min, Asim Kadav, Hans Peter GrafCVPR 2020
