Contrastive Context-Speech Pretraining for Expressive Text-to-Speech Synthesis
Yujia Xiao, Xi Wang, Xu Tan, Lei He, Xinfa Zhu, Sheng Zhao, Tan Lee
摘要
The latest Text-to-Speech (TTS) systems can produce speech with voice quality and naturalness comparable to human speech. Yet the demand for large amount of high-quality data from target speakers remains a significant challenge. Particularly for long-form expressive reading, target speaker's training speech that covers rich contextual information are needed. In this paper a novel design of context-aware speech pre-trained model is developed for expressive TTS based on contrastive learning. The model can be trained with abundant speech data without explicitly labelled speaker identities. It captures the intricate relationship between the speech expression of a spoken sentence and the contextual text information. By incorporating cross-modal text and speech features into the TTS model, it enables the generation of coherent and expressive speech, which is especially beneficial when there is a scarcity of target speaker data. The pre-trained model is evaluated first in the task of Context-Speech retrieval and then as the integral part of a zero-shot TTS system. Experimental results demonstrate that the pretraining framework effectively learns Context-Speech representations and significantly enhances the expressiveness of synthesized speech. Audio demos are available at: https://ccsp2024.github.io/demo/.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 等EMNLP 2025 · 被引用 10 次
- APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechZhicheng Lian, Lizhi Wang, Hua HuangACM MM 2025 · 被引用 1 次
相关 Paper
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang 等ACL 2023 · 被引用 13 次
- UniSpeech: Unified Speech Representation Learning with Labeled and Unlabeled DataChengyi Wang, Yu Wu, Yao Qian, Ken'ichi Kumatani 等ICML 2021 · 被引用 140 次
- A3T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and EditingHe Bai, Renjie Zheng, Jun-Kun Chen, Mingbo Ma 等ICML 2022 · 被引用 64 次
- Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-trainingYifan Yang, Bing Han, Hui Wang, Wei Wang 等ACL 2026 · 被引用 4 次
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin 等ACM MM 2024 · 被引用 15 次
