SpeechTripleNet: End-to-End Disentangled Speech Representation Learning for Content, Timbre and Prosody
Hui Lu, Xixin Wu, Zhiyong Wu, Helen Meng
摘要
Disentangled speech representation learning aims to separate different factors of variation from speech into disjoint representations. This paper focuses on disentangling speech into representations for three factors: spoken content, speaker timbre, and speech prosody. Many previous methods for speech disentanglement have focused on separating spoken content and speaker timbre. However, the lack of explicit modeling of prosodic information leads to degraded speech generation performance and uncontrollable prosody leakage into content and/or speaker representations. While some recent methods have utilized explicit speaker labels or pre-trained models to facilitate triple-factor disentanglement, there are no end-to-end methods to simultaneously disentangle three factors using only unsupervised or self-supervised learning objectives. This paper introduces SpeechTripleNet, an end-to-end method to disentangle speech into representations for content, timbre, and prosody. Based on VAE, SpeechTripleNet restricts the structures of the latent variables and the amount of information captured in them to induce disentanglement. It is a pure unsupervised/self-supervised learning method that only requires speech data and no additional labels. Our qualitative and quantitative results demonstrate that SpeechTripleNet is effective in achieving triple-factor speech disentanglement, as well as controllable speech editing concerning different factors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language ModellingYixuan Zhou, Xiaoyu Qin, Zeyu Jin, Shuoyi Zhou 等ACM MM 2024 · 被引用 10 次
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 等EMNLP 2025 · 被引用 10 次
它引用的顶会 Paper2
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson 等ICML 2020 · 被引用 210 次
相关 Paper
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan 等ICML 2024 · 被引用 341 次
- DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech CodecTao Li, Wenshuo Ge, Zhichao Wang, Zihao Cui 等ACL 2026 · 被引用 1 次
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie 等AAAI 2023 · 被引用 15 次
- Disentangling Voice and Content with Self-Supervision for Speaker RecognitionTianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou LiNeurIPS 2023 · 被引用 53 次
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni 等ICML 2022 · 被引用 157 次
