PortaSpeech: Portable and High-Quality Generative Text-to-Speech
Yi Ren, Jinglin Liu, Zhou Zhao
Abstract
Non-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 [25] and Glow-TTS [8] can synthesize high-quality speech from the given text in parallel. After analyzing two kinds of generative NAR-TTS models (VAE and normalizing flow), we find that: VAE is good at capturing the long-range semantics features (e.g., prosody) even with small model size but suffers from blurry and unnatural results; and normalizing flow is good at reconstructing the frequency bin-wise details but performs poorly when the number of model parameters is limited. Inspired by these observations, to generate diverse speech with natural details and rich prosody using a lightweight architecture, we propose PortaSpeech, a portable and high-quality generative text-to-speech model. Specifically, 1) to model both the prosody and mel-spectrogram details accurately, we adopt a lightweight VAE with an enhanced prior followed by a flow-based post-net with strong conditional inputs as the main architecture. 2) To further compress the model size and memory footprint, we introduce the grouped parameter sharing mechanism to the affine coupling layers in the post-net. 3) To improve the expressiveness of synthesized speech and reduce the dependency on accurate fine-grained alignment between text and speech, we propose a linguistic encoder with mixture alignment combining hard word-level alignment and soft phoneme-level alignment, which explicitly extracts word-level semantic information. Experimental results show that PortaSpeech outperforms other TTS models in both voice quality and prosody modeling in terms of subjective and objective evaluation metrics, and shows only a slight performance degradation when reducing the model parameters to 6.7M (about 4x model size and 3x runtime memory compression ratio compared with FastSpeech 2). Our extensive ablation studies demonstrate that each design in PortaSpeech is effective 3 . * Equal contribution. † Corresponding author 3 Audio samples are available at https://portaspeech.github.io/ . 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5e9ebc2-e11d-42e7-bdcc-fd24feb0fbecCited by top-tier papers11
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
- GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechRongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui et al.NeurIPS 2022 · 99 citations
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song et al.NeurIPS 2022 · 81 citations
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionHa-Yeong Choi, Sang-Hoon Lee, Seong-Whan LeeAAAI 2024 · 66 citations
- GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisZhenhui Ye, Ziyue Jiang, Yi Ren, Jinglin Liu et al.ICLR 2023 · 33 citations
Builds on8
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
- Non-Autoregressive Neural Text-to-SpeechKainan Peng, Wei Ping, Zhao Song, Kexin ZhaoICML 2020 · 118 citations
Related papers
- Revisiting Over-Smoothness in Text to SpeechYi Ren, Xu Tan, Tao Qin, Zhou Zhao et al.ACL 2022
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 42 citations
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian et al.AAAI 2026 · 3 citations
- EfficientTTS: An Efficient and High-Quality Text-to-Speech ArchitectureChenfeng Miao, Shuang Liang, Zhengchen Liu, Minchuan Chen et al.ICML 2021 · 45 citations
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu et al.AAAI 2022 · 21 citations
