StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
Jixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning, Jianhao Ye, Hongbin Zhou, Lei Xie
Abstract
Zero-shot voice conversion (VC) aims to transfer the timbre from the source speaker to an arbitrary unseen speaker while preserving the original linguistic content. Despite recent advancements in zero-shot VC using language model-based or diffusion-based approaches, several challenges remain: 1) current approaches primarily focus on adapting timbre from unseen speakers and are unable to transfer style and timbre to different unseen speakers independently; 2) these approaches often suffer from slower inference speeds due to the autoregressive modeling methods or the need for numerous sampling steps; 3) the quality and similarity of the converted samples are still not fully satisfactory. To address these challenges, we propose a Style controllable zero-shot VC approach named StableVC, which aims to transfer timbre and style from source speech to different unseen target speakers. Specifically, we decompose speech into linguistic content, timbre, and style, and then employ a conditional flow matching module to reconstruct the high-quality mel-spectrogram based on these decomposed features. To effectively capture timbre and style in a zero-shot manner, we introduce a novel dual attention mechanism with an adaptive gate, rather than using conventional feature concatenation. With this non-autoregressive design, StableVC can efficiently capture the intricate timbre and style from different unseen speakers and generate high-quality speech significantly faster than real-time. Experiments demonstrate that our proposed StableVC outperforms state-of-the-art baseline systems in zero-shot VC and achieves flexible control over timbre and style from different unseen speakers. Moreover, StableVC offers approximately 25x and 1.65x faster sampling compared to autoregressive and diffusion-based baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87b4d17a-e5f5-4fba-bf83-1973166c0c58Cited by top-tier papers6
- MusFlow: Multimodal Music Generation via Conditional Flow MatchingJiahao Song, Yuzhao WangACM MM 2025 · 3 citations
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice EnhancingGaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang et al.ACM MM 2025 · 2 citations
- StreamFlow: Streaming Audio Generation from Discrete Tokens via Streaming Flow MatchingHa-Yeong Choi, Sang-Hoon LeeNeurIPS 2025 · 2 citations
- GenSE: Generative Speech Enhancement via Language Models using Hierarchical ModelingJixun Yao, Hexin Liu, Chen Chen, Yuchen Hu et al.ICLR 2025
- Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre ModelingYuguang Yang, Yu Pan, Jixun Yao, Xiang Zhang et al.ACL 2025
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Voicebox: Text-Guided Multilingual Universal Speech Generation at ScaleMatthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer et al.NeurIPS 2023 · 613 citations
Related papers
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource ScenariosBingsong Bai, Yizhong Geng, Fengping Wang, Cong Wang et al.AAAI 2026 · 1 citation
- Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow MatchingJialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li et al.ACL 2025 · 1 citation
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan et al.EMNLP 2024 · 4 citations
- Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised DisentanglementXueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang et al.ICLR 2025
