BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
Yue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Xin He, Qu Yang, Qingxuan Jiang, Fanghua Ye
摘要
The rise of Large Language Models (LLMs) is reshaping multimodel models, with speech synthesis being a prominent application. However, existing approaches often underutilize the linguistic intelligence of these models, typically failing to leverage their powerful instruction-following capabilities. This limitation hinders the model's ability to follow text instructions for controllable Text-to-Speech (TTS). To address this, we propose a new paradigm inspired by operationalism''that decouples instruction understanding from speech generation. We introduce BatonVoice, a framework where an LLM acts as a conductor'', understanding user instructions and generating a textual plan''-- explicit vocal features (e.g., pitch, energy). A separate TTS model, the orchestra'', then generates the speech from these features. To realize this component, we develop BatonTTS, a TTS model trained specifically for this task. Our experiments demonstrate that BatonVoice achieves strong performance in controllable and emotional speech synthesis, outperforming strong open- and closed-source baselines. Notably, our approach enables remarkable zero-shot cross-lingual generalization, accurately applying feature control abilities to languages unseen during post-training. This demonstrates that objectifying speech into textual vocal features can more effectively unlock the linguistic intelligence of LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong 等NeurIPS 2025 · 被引用 181 次
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He 等ICLR 2024 · 被引用 75 次
- Interleaving Reasoning for Better Text-to-Image GenerationWenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao 等ICLR 2026 · 被引用 40 次
- ImageGen-CoT: Enhancing Text-to-Image in-context Learning with Chain-of-Thought ReasoningJiaqi Liao, Zhengyuan Yang, Linjie Li, Dianqi Li 等ICCV 2025 · 被引用 4 次
相关 Paper
- DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech CodecTao Li, Wenshuo Ge, Zhichao Wang, Zihao Cui 等ACL 2026 · 被引用 1 次
- FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language InstructionsDekun Chen, Xueyao Zhang, Yuancheng Wang, Kenan Dai 等ICLR 2026 · 被引用 21 次
- Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask LearnersRongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang 等ACL 2024 · 被引用 6 次
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma 等ACM MM 2025 · 被引用 16 次
- Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of SpeechJianxing Yu, Zihao Gou, Chen Li, Zhisheng Wang 等EMNLP 2025
