Inferring Speaking Styles from Multi-modal Conversational Context by Multi-scale Relational Graph Convolutional Networks
Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, Qiao Tian, Yuping Wang, Yuxuan Wang
Abstract
To support applications of speech-driven interactive systems in various conversational scenarios, text-to-speech (TTS) synthesis needs to understand the conversational context and determine appropriate speaking styles in its synthesized speeches. These speaking styles are influenced by the dependencies between the multi-modal information in the context at both global scale (i.e. utterance level) and local scale (i.e. word level). However, the dependency modeling and speaking style inference at the local scale are largely missing in state-of-the-art TTS systems, resulting in the synthesis of incorrect or improper speaking styles. In this paper, to learn the dependencies in conversations at both global and local scales and to improve the synthesis of speaking styles, we propose a context modeling method which models the dependencies among the multi-modal information in context with multi-scale relational graph convolutional network (MSRGCN). The learnt multi-modal context information at multiple scales is then utilized to infer the global and local speaking styles of the current utterance for speech synthesis. Experiments demonstrate the effectiveness of the proposed approach, and ablation studies reflect the contributions from modeling multi-modal information and multi-scale dependencies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe9e7b5d-b2d2-4374-a8be-f0ff82e8986bCited by top-tier papers3
- Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context ModelingRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.AAAI 2024 · 31 citations
- Generative Expressive Conversational Speech SynthesisRui Liu, Yifan Hu, Yi Ren, Xiang Yin et al.ACM MM 2024 · 15 citations
- UniTalker: Conversational Speech-Visual SynthesisYifan Hu, Rui Liu, Yi Ren, Xiang Yin et al.ACM MM 2025 · 2 citations
Builds on5
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion RecognitionWeizhou Shen, Junqing Chen, Xiaojun Quan, Zhixian XieAAAI 2021 · 251 citations
- Relation-aware Graph Attention Networks with Relational Position Encodings for Emotion Recognition in ConversationsTaichi Ishiwatari, Yuki Yasuda, Taro Miyazaki, Jun GotoEMNLP 2020 · 201 citations
Related papers
- Emotionally Situated Text-to-Speech Synthesis in User-Agent ConversationYuchen Liu, Haoyu Zhang, Shichao Liu, Xiang Yin et al.ACM MM 2023 · 6 citations
- SeDepTTS: Enhancing the Naturalness via Semantic Dependency and Local Convolution for Text-to-Speech SynthesisChenglong Jiang, Ying Gao, Wing W. Y. Ng, Jiyong Zhou et al.AAAI 2023 · 4 citations
- ArtSpeech: Adaptive Text-to-Speech Synthesis with Articulatory RepresentationsZhongxu Wang, Yujia Wang, Mingzhu Li, Hua HuangACM MM 2024 · 2 citations
- CMCU-CSS: Enhancing Naturalness via Commonsense-based Multi-modal Context Understanding in Conversational Speech SynthesisYayue Deng, Jinlong Xue, Fengping Wang, Yingming Gao et al.ACM MM 2023 · 7 citations
- Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech SynthesisSang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim et al.AAAI 2021 · 60 citations
