BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
Yue Wang, Ruotian Ma, Xingyu Chen, Zhengliang Shi, Morunliu Yang, Wanshun Chen, Huang Liu, Jiadi Yao, Xin He, Qu Yang, Qingxuan Jiang, Fanghua Ye
Abstract
The rise of Large Language Models (LLMs) is reshaping multimodel models, with speech synthesis being a prominent application. However, existing approaches often underutilize the linguistic intelligence of these models, typically failing to leverage their powerful instruction-following capabilities. This limitation hinders the model's ability to follow text instructions for controllable Text-to-Speech (TTS). To address this, we propose a new paradigm inspired by operationalism''that decouples instruction understanding from speech generation. We introduce BatonVoice, a framework where an LLM acts as a conductor'', understanding user instructions and generating a textual plan''-- explicit vocal features (e.g., pitch, energy). A separate TTS model, the orchestra'', then generates the speech from these features. To realize this component, we develop BatonTTS, a TTS model trained specifically for this task. Our experiments demonstrate that BatonVoice achieves strong performance in controllable and emotional speech synthesis, outperforming strong open- and closed-source baselines. Notably, our approach enables remarkable zero-shot cross-lingual generalization, accurately applying feature control abilities to languages unseen during post-training. This demonstrates that objectifying speech into textual vocal features can more effectively unlock the linguistic intelligence of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77af4e33-e8e8-4832-b691-b5605409d5edBuilds on6
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He et al.ICLR 2024 · 75 citations
- Interleaving Reasoning for Better Text-to-Image GenerationWenxuan Huang, Shuang Chen, Zheyong Xie, Shaosheng Cao et al.ICLR 2026 · 40 citations
- ImageGen-CoT: Enhancing Text-to-Image in-context Learning with Chain-of-Thought ReasoningJiaqi Liao, Zhengyuan Yang, Linjie Li, Dianqi Li et al.ICCV 2025 · 4 citations
Related papers
- DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech CodecTao Li, Wenshuo Ge, Zhichao Wang, Zihao Cui et al.ACL 2026 · 1 citation
- FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language InstructionsDekun Chen, Xueyao Zhang, Yuancheng Wang, Kenan Dai et al.ICLR 2026 · 21 citations
- Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask LearnersRongjie Huang, Chunlei Zhang, Yongqi Wang, Dongchao Yang et al.ACL 2024 · 6 citations
- EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingGuanrou Yang, Chen Yang, Qian Chen, Ziyang Ma et al.ACM MM 2025 · 16 citations
- Eliciting Implicit Acoustic Styles from Open-domain Instructions to Facilitate Fine-grained Controllable Generation of SpeechJianxing Yu, Zihao Gou, Chen Li, Zhisheng Wang et al.EMNLP 2025
