EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
Haozhe Chen, Run Chen, Julia Hirschberg
摘要
While recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity. We propose EmoKnob, a framework that allows fine-grained emotion control in speech synthesis with fewshot demonstrative samples of arbitrary emotion. Our framework leverages the expressive speaker representation space made possible by recent advances in foundation voice cloning models. Based on the few-shot capability of our emotion control framework, we propose two methods to apply emotion control on emotions described by open-ended text, enabling an intuitive interface for controlling a diverse array of nuanced emotions. To facilitate a more systematic emotional speech synthesis field, we introduce a set of evaluation metrics designed to rigorously assess the faithfulness and recognizability of emotion control frameworks. Through objective and subjective evaluations, we show that our emotion control framework effectively embeds emotions into speech and surpasses emotion expressiveness of commercial TTS services.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cross-Modal Emotion Transfer for Emotion Editing in Talking Face VideoChanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung 等CVPR 2026 · 被引用 1 次
- ES4R: Speech Encoding Based on Prepositive Affective Modeling for Empathetic Response GenerationZhuoyue Gao, Xiaohui Wang, Xiaocui Yang, Wen Zhang 等ACL 2026
- TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech SynthesisQifan Liang, Yuansen Liu, Ruixin Wei, Nan Lu 等ACL 2026
它引用的顶会 Paper2
相关 Paper
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma 等ACM MM 2025 · 被引用 16 次
- Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisTianrui Wang, Haoyu Wang, Meng Ge, Cheng Gong 等NeurIPS 2025 · 被引用 8 次
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang 等EMNLP 2025 · 被引用 10 次
- High-Fidelity Generalized Emotional Talking Face Generation with Multi-Modal Emotion Space LearningChao Xu, Junwei Zhu, Jiangning Zhang, Yue Han 等CVPR 2023
- CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation SteeringSiyi Wang, Shihong Tan, Siyi Liu, Hong Jia 等ICML 2026
