P-Flow: A Fast and Data-Efficient Zero-Shot TTS through Speech Prompting
Sungwon Kim, Kevin J. Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T. Desta, Rafael Valle, Sungroh Yoon, Bryan Catanzaro
Abstract
While recent large-scale neural codec language models have shown significant improvement in zero-shot TTS by training on thousands of hours of data, they suffer from drawbacks such as a lack of robustness, slow sampling speed similar to previous autoregressive TTS methods, and reliance on pre-trained neural codec representations. Our work proposes P-Flow, a fast and data-efficient zero-shot TTS model that uses speech prompts for speaker adaptation. P-Flow comprises a speech-prompted text encoder for speaker adaptation and a flow matching generative decoder for high-quality and fast speech synthesis. Our speech-prompted text encoder uses speech prompts and text input to generate speaker-conditional text representation. The flow matching generative decoder uses the speaker-conditional output to synthesize high-quality personalized speech significantly faster than in real-time. Unlike the neural codec language models, we specifically train P-Flow on LibriTTS dataset using a continuous mel-representation. Through our training method using continuous speech prompts, P-Flow matches the speaker similarity performance of the large-scale zero-shot TTS models with two orders of magnitude less training data and has more than 20 × faster sampling speed. Our results show that P-Flow has better pronunciation and is preferred in human likeness and speaker similarity to its recent state-of-the-art counterparts, thus defining P-Flow as an attractive and desirable alternative. We provide audio samples on our demo page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd529488-e45d-4776-a34c-aca4e32df3a3Cited by top-tier papers13
- CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker ConversationsLeying Zhang, Yao Qian, Long Zhou, Shujie Liu et al.NeurIPS 2024 · 31 citations
- Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech DeepfakesKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang et al.AAAI 2025 · 5 citations
- FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice EnhancingGaoxiang Cong, Liang Li, Jiadong Pan, Zhedong Zhang et al.ACM MM 2025 · 2 citations
- PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform GenerationSang-Hoon Lee, Ha-Yeong Choi, Seong-Whan LeeICLR 2025
- DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific FactorsKeon Lee, Dong Won Kim, Jaehyeon Kim, Seungjun Chung et al.ICLR 2025
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
Related papers
- Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech SynthesisYifan Yang, Shujie Liu, Jinyu Li, Yuxuan Hu et al.ACM MM 2025 · 1 citation
- FlashSpeech: Efficient Zero-Shot Speech SynthesisZhen Ye, Zeqian Ju, Haohe Liu, Xu Tan et al.ACM MM 2024 · 10 citations
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsYinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler et al.NeurIPS 2023 · 324 citations
- MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-SpeechShengpeng Ji, Ziyue Jiang, Hanting Wang, Jialong Zuo et al.ACL 2024 · 1 citation
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
