HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis
Sang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song, Min-Jae Hwang, Seong-Whan Lee
摘要
This paper presents HierSpeech, a high-quality end-to-end text-to-speech (TTS) system based on a hierarchical conditional variational autoencoder (VAE) utilizing self-supervised speech representations. Recently, single-stage TTS systems, which directly generate raw speech waveform from text, have been getting interest thanks to their ability in generating high-quality audio within a fully end-to-end training pipeline. However, there is still a room for improvement in the conventional TTS systems. Since it is challenging to infer both the linguistic and acoustic attributes from the text directly, missing the details of attributes, specifically linguistic information, is inevitable, which results in mispronunciation and over-smoothing problem in their synthetic speech. To address the aforementioned problem, we leverage self-supervised speech representations as additional linguistic representations to bridge an information gap between text and speech. Then, the hierarchical conditional VAE is adopted to connect these representations and to learn each attribute hierarchically by improving the linguistic capability in latent representations. Compared with the state-of-the-art TTS system, HierSpeech achieves +0.303 comparative mean opinion score, and reduces the phoneme error rate of synthesized speech from 9.16% to 5.78% on the VCTK dataset. Furthermore, we extend our model to HierSpeech-U, an untranscribed text-to-speech system. Specifically, HierSpeech-U can adapt to a novel speaker by utilizing self-supervised speech representations without text transcripts. The experimental results reveal that our method outperforms publicly available TTS models, and show the effectiveness of speaker adaptation with untranscribed speech. * Corresponding author 36th Conference on Neural Information Processing Systems (NeurIPS 2022). text sequence, and the vocoder (Oord et al., 2016) converts the acoustic features into raw waveforms consecutively. However, previous TTS models are subject to two limitations: 1) although speech consists of various attributes (e.g., pronunciation, rhythm, intonation, and timbre) (Qian et al., 2020; Choi et al., 2021) , most previous models synthesize acoustic features from the text sequence at once (Ren et al., 2019) , which exacerbates the one-to-many mapping problem; and 2) in the two-stage pipeline, each part of the TTS system should be trained independently, which results in the degradation of the audio quality (Ren et al., 2021a,b; Lee et al., 2021b). Recently, single-stage end-to-end TTS models, which directly generate a raw waveform from text, successfully reduce these limitations of the two-stage pipeline. For instance, VITS (Kim et al., 2021) adopts variational inference augmented with the normalizing flow (Kim et al., 2020) and adversarial training (Kong et al., 2020) to improve the expressiveness of the model, which can learn rich representations from speech data and synthesize waveforms directly from the text. However, despite efforts to reduce the information gap between text and speech, these models are subject to speech mispronunciation and over-smoothing problems. In the process of synthesizing speech, they still generate all the acoustic attributes from text sequence at the same time. Therefore, missing the details of some attributes between text and speech, specifically linguistic information, is inevitable. To bridge the information gap between text and speech, we adopt self-supervised speech representations as additional linguistic representations. Trained with large-scale speech dataset, these representations can learn useful information without using labeled data. Previous studies (Shah et al., 2021; Choi et al., 2021) also reveal that the representations from the pre-trained model contain rich information trained from a large-scale speech dataset. In particular, the representations from the middle layer of the pre-trained model contain rich linguistic information which has a characteristic of pronunciation. As a result, it has been successfully utilized for various speech tasks such as speech recognition (Baevski et al., 2020 , 2021), voice conversion (Choi et al., 2021; Lee et al., 2021a), and speech resynthesis (Polyak et al., 2021) . However, these useful representations have not yet received significant attention in TTS systems due to the difficulty to utilize in generative model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsYinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler 等NeurIPS 2023 · 被引用 324 次
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionHa-Yeong Choi, Sang-Hoon Lee, Seong-Whan LeeAAAI 2024 · 被引用 66 次
- Disentangling Voice and Content with Self-Supervision for Speaker RecognitionTianchi Liu, Kong Aik Lee, Qiongqiong Wang, Haizhou LiNeurIPS 2023 · 被引用 53 次
- Let There Be Sound: Reconstructing High Quality Speech from Silent VideosJi-Hoon Kim, Jaehun Kim, Joon Son ChungAAAI 2024 · 被引用 14 次
- Multi-View Collaborative Learning Network for Speech Deepfake DetectionKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yifang Guo 等AAAI 2025 · 被引用 8 次
它引用的顶会 Paper13
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 被引用 1,141 次
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
相关 Paper
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 被引用 42 次
- T2V2: A Unified Non-Autoregressive Model for Speech Recognition and Synthesis via Multitask LearningNabarun Goswami, Hanqin Wang, Tatsuya HaradaICLR 2025
- KALL-E: Autoregressive Speech Synthesis with Next-Distribution PredictionKangxiang Xia, Xinfa Zhu, Jixun Yao, Wenjie Tian 等AAAI 2026 · 被引用 3 次
- Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech SynthesisYixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li 等ICLR 2026
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng 等ICLR 2025
