FPETS: Fully Parallel End-to-End Text-to-Speech System
Dabiao Ma, Zhiba Su, Wenxuan Wang, Yuhao Lu
Abstract
End-to-end Text-to-speech (TTS) system can greatly improve the quality of synthesised speech. But it usually suffers form high time latency due to its auto-regressive structure. And the synthesised speech may also suffer from some error modes, e.g. repeated words, mispronunciations, and skipped words. In this paper, we propose a novel non-autoregressive, fully parallel end-to-end TTS system (FPETS). It utilizes a new alignment model and the recently proposed U-shape convolutional structure, UFANS. Different from RNN, UFANS can capture long term information in a fully parallel manner. Trainable position encoding and two-step training strategy are used for learning better alignments. Experimental results show FPETS utilizes the power of parallel computation and reaches a significant speed up of inference compared with state-of-the-art end-to-end TTS systems. More specifically, FPETS is 600X faster than Tacotron2, 50X faster than DCTTS and 10X faster than Deep Voice3. And FPETS can generates audios with equal or better quality and fewer errors comparing with other system. As far as we know, FPETS is the first end-to-end TTS system which is fully parallel.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58e83d9c-0479-4c4f-a392-cc57cb0a5d74Cited by top-tier papers3
- MTTM: Metamorphic Testing for Textual Content Moderation SoftwareWenxuan Wang, Jen-tse Huang, Weibin Wu, Jianping Zhang et al.ICSE 2023 · 23 citations
- An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareWenxuan Wang, Jingyuan Huang, Jen-tse Huang, Chang Chen et al.ASE 2023 · 7 citations
- Metamorphic Testing for Audio Content Moderation SoftwareWenxuan Wang, Yongjiang Wu, Junyuan Zhang, Shuqing Li et al.ASE 2025
Related papers
- EfficientTTS: An Efficient and High-Quality Text-to-Speech ArchitectureChenfeng Miao, Shuang Liang, Zhengchen Liu, Minchuan Chen et al.ICML 2021 · 45 citations
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 42 citations
- Non-Autoregressive Neural Text-to-SpeechKainan Peng, Wei Ping, Zhao Song, Kexin ZhaoICML 2020 · 118 citations
