Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
Tianrui Wang, Haoyu Wang, Meng Ge, Cheng Gong, Chunyu Qiang, Ziyang Ma, Zikang Huang, Guanrou Yang, Xiaobao Wang, Eng-Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang
Abstract
While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions. Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4428533-248e-4f3b-bbaa-d5aefb2df2bcBuilds on8
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan et al.ICML 2024 · 341 citations
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma et al.ACM MM 2025 · 16 citations
Related papers
- TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech SynthesisQifan Liang, Yuansen Liu, Ruixin Wei, Nan Lu et al.ACL 2026
- Semi-Supervised Generative Modeling for Controllable Speech SynthesisRaza Habib, Soroosh Mariooryad, Matt Shannon, Eric Battenberg et al.ICLR 2020 · 48 citations
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He et al.ACM MM 2024 · 3 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
- EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion ControlHaozhe Chen, Run Chen, Julia HirschbergEMNLP 2024 · 4 citations
