CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
Siyi Wang, Shihong Tan, Siyi Liu, Hong Jia, Gongping Huang, James Bailey, Ting Dang
摘要
Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech (TTS) systems enforce a single utterance-level emotion, collapsing affective diversity and suppressing mixed or text–emotion–misaligned expression. While activation steering via latent direction vectors offers a promising solution, it remains unclear whether emotion representations are linearly steerable in TTS, where steering should be applied within hybrid TTS architectures, and how such complex emotion behaviors should be evaluated. This paper presents the first systematic analysis of activation steering for emotional control in hybrid TTS models, introducing a quantitative, controllable steering framework, and multi-rater evaluation protocols that enable composable mixed-emotion synthesis and reliable text–emotion mismatch synthesis. Our results demonstrate, for the first time, that emotional prosody and expressive variability are primarily synthesized by the TTS language module instead of the flow-matching module, and also provide a lightweight steering approach for generating natural, human-like emotional speech.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 被引用 461 次
- StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language ModelsYinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler 等NeurIPS 2023 · 被引用 324 次
- Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference OptimizationYuanpu Cao, Tianrong Zhang, Bochuan Cao, Ziyi Yin 等NeurIPS 2024 · 被引用 135 次
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou 等AAAI 2026 · 被引用 63 次
相关 Paper
- Sparse Autoencoders for Interpretable Emotion Control in Text-to-SpeechHongfei Du, Jiacheng Shi, Sidi Lu, Gang Zhou 等ICML 2026
- AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-SpeechBin Kang, Shaoguo Wen, Yang Fan, Shunlong Wu 等ICML 2026
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 被引用 53 次
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma 等ACM MM 2025 · 被引用 16 次
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 等EMNLP 2025 · 被引用 10 次
