Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech
Yoonhyung Lee, Joongbo Shin, Kyomin Jung
摘要
Although early text-to-speech (TTS) models such as Tacotron 2 have succeeded in generating human-like speech, their autoregressive architectures have several limitations: (1) They require a lot of time to generate a mel-spectrogram consisting of hundreds of steps. (2) The autoregressive speech generation lacks robustness due to its error propagation property. In this paper, we propose a novel non-autoregressive TTS model called BVAE-TTS, which eliminates the architectural limitations and generates a mel-spectrogram in parallel. BVAE-TTS adopts a bidirectional-inference variational autoencoder (BVAE) that learns hierarchical latent representations using both bottom-up and top-down paths to increase its expressiveness. To apply BVAE to TTS, we design our model to utilize text information via an attention mechanism. By using attention maps that BVAE-TTS generates, we train a duration predictor so that the model uses the predicted duration of each phoneme at inference. In experiments conducted on LJSpeech dataset, we show that our model generates a mel-spectrogram 27 times faster than Tacotron 2 with similar speech quality. Furthermore, our BVAE-TTS outperforms Glow-TTS, which is one of the state-of-the-art non-autoregressive TTS models, in terms of both speech quality and inference speed while having 58% fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier GuidanceHeeseung Kim, Sungwon Kim, Sungroh YoonICML 2022 · 被引用 133 次
- Translatotron 2: High-quality direct speech-to-speech translation with voice preservationYe Jia, Michelle Tadmor Ramanovich, Tal Remez, Roi PomerantzICML 2022 · 被引用 107 次
- PortaSpeech: Portable and High-Quality Generative Text-to-SpeechYi Ren, Jinglin Liu, Zhou ZhaoNeurIPS 2021 · 被引用 97 次
- HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech SynthesisSang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song 等NeurIPS 2022 · 被引用 81 次
它引用的顶会 Paper4
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- Non-Autoregressive Neural Text-to-SpeechKainan Peng, Wei Ping, Zhao Song, Kexin ZhaoICML 2020 · 被引用 118 次
相关 Paper
- EfficientTTS: An Efficient and High-Quality Text-to-Speech ArchitectureChenfeng Miao, Shuang Liang, Zhengchen Liu, Minchuan Chen 等ICML 2021 · 被引用 45 次
- FPETS: Fully Parallel End-to-End Text-to-Speech SystemDabiao Ma, Zhiba Su, Wenxuan Wang, Yuhao LuAAAI 2020 · 被引用 7 次
- Revisiting Over-Smoothness in Text to SpeechYi Ren, Xu Tan, Tao Qin, Zhou Zhao 等ACL 2022
- Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech SynthesisRafael Valle, Kevin J. Shih, Ryan Prenger, Bryan CatanzaroICLR 2021 · 被引用 133 次
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu 等AAAI 2022 · 被引用 21 次
