Non-Autoregressive Neural Text-to-Speech
Kainan Peng, Wei Ping, Zhao Song, Kexin Zhao
Abstract
In this work, we propose ParaNet, a non-autoregressive seq2seq model that converts text to spectrogram. It is fully convolutional and brings 46.7 times speed-up over the lightweight Deep Voice 3 at synthesis, while obtaining reasonably good speech quality. ParaNet also produces stable alignment between text and speech on the challenging test sentences by iteratively improving the attention in a layer-by-layer manner. Furthermore, we build the parallel text-to-speech system and test various parallel neural vocoders, which can synthesize speech from text through a single feed-forward pass. We also explore a novel VAE-based approach to train the inverse autoregressive flow (IAF) based parallel vocoder from scratch, which avoids the need for distillation from a separately trained WaveNet as previous work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a88fc5e7-9896-41b1-a91b-1912f6aff6ebCited by top-tier papers19
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- It's Raw! Audio Generation with State-Space ModelsKaran Goel, Albert Gu, Chris Donahue, Christopher RéICML 2022 · 257 citations
- Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech GenerationDongchan Min, Dong Bok Lee, Eunho Yang, Sung Ju HwangICML 2021 · 218 citations
Builds on2
Related papers
- FPETS: Fully Parallel End-to-End Text-to-Speech SystemDabiao Ma, Zhiba Su, Wenxuan Wang, Yuhao LuAAAI 2020 · 7 citations
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 42 citations
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 9 citations
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech SynthesisRafael Valle, Kevin J. Shih, Ryan Prenger, Bryan CatanzaroICLR 2021 · 133 citations
- Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow MatchingYongqi Wang, Wenxiang Guo, Rongjie Huang, Jiawei Huang et al.NeurIPS 2024 · 73 citations
