Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling
Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, Haizhou Li
Abstract
Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problems due to the scarcity of emotional conversational datasets and the difficulty of stateful emotion modeling. In this paper, we propose a novel emotional CSS model, termed ECSS, that includes two main components: 1) to enhance emotion understanding, we introduce a heterogeneous graph-based emotional context modeling mechanism, which takes the multi-source dialogue history as input to model the dialogue context and learn the emotion cues from the context; 2) to achieve emotion rendering, we employ a contrastive learning-based emotion renderer module to infer the accurate emotion style for the target utterance. To address the issue of data scarcity, we meticulously create emotional labels in terms of category and intensity, and annotate additional emotional information on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in understanding and rendering emotions. These evaluations also underscore the importance of comprehensive emotional annotations. Code and audio samples can be found at: https://github.com/walker-hyf/ECSS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92359ce6-3f3b-4648-b7e3-fa3ecb42b26aCited by top-tier papers5
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
- BIG-FUSION: Brain-Inspired Global-Local Context Fusion Framework for Multimodal Emotion Recognition in ConversationsYusong Wang, Xuanye Fang, Huifeng Yin, Dongyuan Li et al.AAAI 2025 · 11 citations
- Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-SpeechRui Liu, Shuwei He, Yifan Hu, Haizhou LiAAAI 2025 · 8 citations
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin et al.ACM MM 2025 · 2 citations
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
Builds on6
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-SpeechRongjie Huang, Yi Ren, Jinglin Liu, Chenye Cui et al.NeurIPS 2022 · 99 citations
Related papers
- CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisYayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao et al.ACM MM 2023 · 7 citations
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Infusing Multi-Source Knowledge with Heterogeneous Graph Neural Network for Emotional Conversation GenerationYunlong Liang, Fandong Meng, Ying Zhang, Yufeng Chen et al.AAAI 2021 · 62 citations
- Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in ConversationZijian Yi, Ziming Zhao, Zhishu Shen, Tiehua ZhangACM MM 2024 · 32 citations
