Generative Expressive Conversational Speech Synthesis
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, Haizhou Li
摘要
Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. To address the above issues, we propose a novel generative expressive CSS system, termed GPT-Talker.We transform the multimodal information of the multi-turn dialogue history into discrete token sequences and seamlessly integrate them to form a comprehensive user-agent dialogue context. Leveraging the power of GPT, we predict the token sequence, that includes both semantic and style knowledge, of response for the agent. After that, the expressive conversational speech is synthesized by the conversation-enriched VITS to deliver feedback to the user.Furthermore, we propose a large-scale Natural CSS Dataset called NCSSD, that includes both naturally recorded conversational speech in improvised styles and dialogues extracted from TV shows. It encompasses both Chinese and English languages, with a total duration of 236 hours. We conducted comprehensive experiments on the reliability of the NCSSD and the effectiveness of our GPT-Talker. Both subjective and objective evaluations demonstrate that our model outperforms other state-of-the-art CSS systems significantly in terms of naturalness and expressiveness. The Code, Dataset, and Pre-trained Model are available at: https://github.com/AI-S2-Lab/GPT-Talker.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language InstructionsDekun Chen, Xueyao Zhang, Yuancheng Wang, Kenan Dai 等ICLR 2026 · 被引用 21 次
- CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow MatchingLeying Zhang, Yao Qian, Xiaofei Wang, Manthan Thakker 等NeurIPS 2025 · 被引用 16 次
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan 等ICLR 2026 · 被引用 7 次
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin 等ACM MM 2025 · 被引用 2 次
- APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechZhicheng Lian, Lizhi Wang, Hua HuangACM MM 2025 · 被引用 1 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue DatabaseJinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu 等ACL 2022 · 被引用 88 次
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He 等ICLR 2024 · 被引用 75 次
相关 Paper
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin 等ACM MM 2023 · 被引用 6 次
- Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context ModelingRui Liu, Yifan Hu, Yi Ren, Xiang Yin 等AAAI 2024 · 被引用 31 次
- Contrastive Context-Speech Pretraining for Expressive Text-to-Speech SynthesisYujia Xiao, Xi Wang, Xu Tan, Lei He 等ACM MM 2024 · 被引用 3 次
- DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and ActionsChuanqi Cheng, Hongda Sun, Bo Du, Shuo Shang 等ACL 2025
- Scaling Rich Style-Prompted Text-to-Speech DatasetsAnuj Diwan, Zhisheng Zheng, David Harwath, Eunsol ChoiEMNLP 2025 · 被引用 2 次
