CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech Synthesis
Yayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao, Ya Li
Abstract
Conversational Speech Synthesis (CSS) aims to produce speech appropriate for oral communication. However, the complexity of context dependency modeling poses significant challenges in the field of CSS, especially the mutual psychological influence between interlocutors. Previous studies have verified that prior commonsense knowledge helps machines understand subtle psychological information (e.g., feelings and intentions) in spontaneous oral dialogues. Therefore, to enhance context understanding and improve the naturalness of synthesized speech, we propose a novel conversational speech synthesis system (CMCU-CSS) that incorporates the Commonsense-based Multi-modal Context Understanding (CMCU) module to model the dynamic emotional interaction among interlocutors. Specifically, we first utilize three implicit states (intent state, internal state and external state) in CMCU to model the context dependency between inter/intra speakers with the help of commonsense knowledge. Furthermore, we infer emotion vectors from the fusion of these implicit states and multi-modal features to enhance the emotion discriminability of synthesized speech. This is the first attempt to combine commonsense knowledge with conversational speech synthesis, and its effect in terms of emotion discriminability of synthetic speech is evaluated by emotion recognition in conversation task. The results of subjective and objective evaluations demonstrate that the CMCU-CSS model achieves more natural speech with context-appropriate emotion and is equipped with the best emotion discriminability, surpassing that of other conversational speech synthesis models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8d0793ad-882f-4fec-b57a-ff2af328c24aCited by top-tier papers3
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
- APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechZhicheng Lian, Lizhi Wang, Hua HuangACM MM 2025 · 1 citation
- Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction LearningRui Liu, Yuan Zhao, Zhenqi JiaAAAI 2026
Related papers
- Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context ModelingRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.AAAI 2024 · 31 citations
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin et al.ACM MM 2025 · 2 citations
- Knowledge Bridging for Empathetic Dialogue GenerationQintong Li, Piji Li, Zhaochun Ren, Pengjie Ren et al.AAAI 2022 · 128 citations
- From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed DialoguesShivani Kumar, Ramaneswaran S., Md. Shad Akhtar, Tanmoy ChakrabortyEMNLP 2023 · 5 citations
