Revisiting Over-Smoothness in Text to Speech
Yi Ren, Xu Tan, Tao Qin, Zhou Zhao, Tie-Yan Liu
Abstract
Non-autoregressive text to speech (NAR-TTS) models have attracted much attention from both academia and industry due to their fast generation speed. One limitation of NAR-TTS models is that they ignore the correlation in time and frequency domains while generating speech mel-spectrograms, and thus cause blurry and over-smoothed results. In this work, we revisit this over-smoothing problem from a novel perspective: the degree of over-smoothness is determined by the gap between the complexity of data distributions and the capability of modeling methods. Both simplifying data distributions and improving modeling methods can alleviate the problem. Accordingly, we first study methods reducing the complexity of data distributions. Then we conduct a comprehensive study on NAR-TTS models that use some advanced modeling methods. Based on these studies, we find that 1) methods that provide additional condition inputs reduce the complexity of data distributions to model, thus alleviating the over-smoothing problem and achieving better voice quality. 2) Among advanced modeling methods, Laplacian mixture loss performs well at modeling multimodal distributions and enjoys its simplicity, while GAN and Glow achieve the best voice quality while suffering from increased training or model complexity. 3) The two categories of methods can be combined to further alleviate the over-smoothness and improve the voice quality. 4) Our experiments on the multi-speaker dataset lead to similar conclusions as above and providing more variance information can reduce the difficulty of modeling the target data distribution and alleviate the requirements for model capacity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c494d04-5274-4674-8a77-bba072de5810Cited by top-tier papers7
- ProDiff: Progressive Fast Diffusion Model for High-Quality Text-to-SpeechRongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu et al.ACM MM 2022 · 182 citations
- MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech SynthesisWenhao Guan, Yishuang Li, Tao Li, Hukai Huang et al.AAAI 2024 · 25 citations
- FlashSpeech: Efficient Zero-Shot Speech SynthesisZhen Ye, Zeqian Ju, Haohe Liu, Xu Tan et al.ACM MM 2024 · 10 citations
- Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioShao-En Weng, Hong-Han Shuai, Wen-Huang ChengAAAI 2023 · 4 citations
- SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech SynthesisChenglong Jiang, Ying Gao, Wing W. Y. Ng, Jiyong Zhou et al.AAAI 2023 · 4 citations
Builds on5
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 663 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- High Fidelity Speech Synthesis with Adversarial NetworksMikolaj Binkowski, Jeff Donahue, Sander Dieleman, Aidan Clark et al.ICLR 2020 · 263 citations
- Non-Autoregressive Neural Text-to-SpeechKainan Peng, Wei Ping, Zhao Song, Kexin ZhaoICML 2020 · 118 citations
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim et al.AAAI 2021 · 60 citations
Related papers
- PortaSpeech: Portable and High-Quality Generative Text-to-SpeechYi Ren, Jinglin Liu, Zhou ZhaoNeurIPS 2021 · 97 citations
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 42 citations
- Flow-Based Unconstrained Lip to Speech GenerationJinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu et al.AAAI 2022 · 21 citations
- A Study of Non-autoregressive Model for Sequence GenerationYi Ren, Jinglin Liu, Xu Tan, Zhou Zhao et al.ACL 2020 · 58 citations
- Chunked Autoregressive GAN for Conditional Waveform SynthesisMax Morrison, Rithesh Kumar, Kundan Kumar, Prem Seetharaman et al.ICLR 2022 · 91 citations
