Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, Haizhou Li
摘要
Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problems due to the scarcity of emotional conversational datasets and the difficulty of stateful emotion modeling. In this paper, we propose a novel emotional CSS model, termed ECSS, that includes two main components: 1) to enhance emotion understanding, we introduce a heterogeneous graph-based emotional context modeling mechanism, which takes the multi-source dialogue history as input to model the dialogue context and learn the emotion cues from the context; 2) to achieve emotion rendering, we employ a contrastive learning-based emotion renderer module to infer the accurate emotion style for the target utterance. To address the issue of data scarcity, we meticulously create emotional labels in terms of category and intensity, and annotate additional emotional information on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in understanding and rendering emotions. These evaluations also underscore the importance of comprehensive emotional annotations. Code and audio samples can be found at: https://github.com/walker-hyf/ECSS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin 等ACM MM 2024 · 被引用 15 次
- BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in ConversationsYusong Wang, Xuanye Fang, Huifeng Yin, Dongyuan Li 等AAAI 2025 · 被引用 11 次
- Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-SpeechRui Liu, Shuwei He, Yifan Hu, Haizhou LiAAAI 2025 · 被引用 8 次
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin 等ACM MM 2025 · 被引用 2 次
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
它引用的顶会 Paper6
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechRongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui 等NeurIPS 2022 · 被引用 99 次
相关 Paper
- CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisYayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao 等ACM MM 2023 · 被引用 7 次
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin 等ACM MM 2023 · 被引用 6 次
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 被引用 46 次
- Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation GenerationYunlong Liang, Fandong Meng, Ying Zhang, Yufeng Chen 等AAAI 2021 · 被引用 62 次
- Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in ConversationZijian Yi, Ziming Zhao, Zhishu Shen, Tiehua ZhangACM MM 2024 · 被引用 32 次
