Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning
Siyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao, Zhe Gan, Lawrence Carin
Abstract
Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel training data and pre-known speakers. However, zero-shot voice style transfer, which learns from non-parallel data and generates voices for previously unseen speakers, remains a challenging problem. We propose a novel zero-shot voice transfer method via disentangled representation learning. The proposed method first encodes speaker-related style and voice content of each input voice into separated low-dimensional embedding spaces, and then transfers to a new voice by combining the source content embedding and target style embedding through a decoder. With information-theoretic guidance, the style and content embedding spaces are representative and (ideally) independent of each other. On real-world VCTK datasets, our method outperforms other baselines and obtains state-of-the-art results in terms of transfer accuracy and voice naturalness for voice style transfer experiments under both many-to-many and zero-shot setups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66517cd5-d0bd-4591-8162-efed61b63dc3Cited by top-tier papers13
- Semantic Feature Extraction for Generalized Zero-Shot LearningJunhan Kim, Kyuhong Shim, Byonghyo ShimAAAI 2022 · 46 citations
- VoiceMixer: Adversarial Voice Style MixupSang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, Seong-Whan LeeNeurIPS 2021 · 46 citations
- Learning the Beauty in Songs: Neural Singing Voice BeautifierJinglin Liu, Chengxi Li, Yi Ren, Zhiying Zhu et al.ACL 2022 · 25 citations
- Retriever: Learning Content-Style Representation as a Token-Level Bipartite GraphDacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang et al.ICLR 2022 · 13 citations
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingJixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning et al.AAAI 2025 · 13 citations
Builds on2
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
- Improving Disentangled Text Representation Learning with Information-Theoretic GuidancePengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon et al.ACL 2020 · 66 citations
Related papers
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai et al.ACM MM 2021 · 15 citations
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementXueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang et al.ICLR 2025
- Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice ConversionYan Rong, Li LiuAAAI 2025 · 11 citations
- HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosBingsong Bai, Yizhong Geng, Fengping Wang, Cong Wang et al.AAAI 2026 · 1 citation
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
