TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
Qifan Liang, Yuansen Liu, Ruixin Wei, Nan Lu, Junchuan Zhao, Ye Wang
Abstract
While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public datasets or complex multi-stage training. In this paper, we propose TED-TTS, a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. Specifically, we propose a segment-aware emotion conditioning strategy that combines causal masking with monotonic stream alignment filtering to isolate emotion conditioning and schedule mask transitions, enabling smooth intra-utterance emotion shifts while preserving global semantic coherence. Based on this, we further propose a segment-aware duration steering strategy to combine local duration embedding steering with global EOS logit modulation, allowing local duration adjustment while ensuring globally consistent termination. To eliminate the need for segment-level manual prompt engineering, we construct a 30,000-sample multi-emotion and duration-annotated text dataset to enable LLM-based automatic prompt construction. Extensive experiments demonstrate that our training-free method not only achieves state-of-the-art intra-utterance consistency in multi-emotion and duration control, but also maintains baseline-level speech quality of the underlying TTS model. Code and audio samples are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6571d729-fed5-4e1b-a3f6-0b84245f8569Builds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma et al.ACM MM 2025 · 16 citations
- VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language ModellingYixuan Zhou, Xiaoyu Qin, Zeyu Jin, Shuoyi Zhou et al.ACM MM 2024 · 10 citations
- EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion ControlHaozhe Chen, Run Chen, Julia HirschbergEMNLP 2024 · 4 citations
Related papers
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisTianrui Wang, Haoyu Wang, Meng Ge, Cheng Gong et al.NeurIPS 2025 · 8 citations
- MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec TransformerYuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng et al.ICLR 2025
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan et al.ICLR 2026 · 7 citations
- Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free GuidanceShehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova et al.EMNLP 2025 · 2 citations
- Contrastive Context-Speech Pretraining for Expressive Text-to-Speech SynthesisYujia Xiao, Xi Wang, Xu Tan, Lei He et al.ACM MM 2024 · 3 citations
