Unsupervised Speech Decomposition via Triple Information Bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, David D. Cox
Abstract
Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm. Obtaining disentangled representations of these components is useful in many speech analysis and generation applications. Recently, state-of-the-art voice conversion systems have led to speech representations that can disentangle speaker-dependent and independent information. However, these systems can only disentangle timbre, while information about pitch, rhythm and content is still mixed together. Further disentangling the remaining speech components is an under-determined problem in the absence of explicit annotations for each component, which are difficult and expensive to obtain. In this paper, we propose SPEECHSPLIT, which can blindly decompose speech into its four components by introducing three carefully designed information bottlenecks. SPEECHSPLIT is among the first algorithms that can separately perform style transfer on timbre, pitch and rhythm without text labels. Our code is publicly available at https://github.com/auspicious3000/ SpeechSplit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8db79cf1-73be-4464-87ec-12c291af20b5Cited by top-tier papers28
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan et al.ICML 2024 · 341 citations
- Graph Information Bottleneck for Subgraph RecognitionJunchi Yu, Tingyang Xu, Yu Rong, Yatao Bian et al.ICLR 2021 · 200 citations
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language ModelsXin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou et al.ICLR 2024 · 126 citations
- Chunked Autoregressive GAN for Conditional Waveform SynthesisMax Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman et al.ICLR 2022 · 91 citations
Related papers
- Global Prosody Style Transfer Without Text TranscriptionsKaizhi Qian, Yang Zhang, Shiyu Chang, Jinjun Xiong et al.ICML 2021 · 25 citations
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- VoiceMixer: Adversarial Voice Style MixupSang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, Seong-Whan LeeNeurIPS 2021 · 46 citations
- PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice ConversionYimin Deng, Huaizhen Tang, Xulong Zhang, Jianzong Wang et al.ACM MM 2023 · 15 citations
- SpeechTripleNet: End-to-End Disentangled Speech Representation Learning for Content, Timbre and ProsodyHui Lu, Xixin Wu, Zhiyong Wu, Helen MengACM MM 2023 · 5 citations
