UniVocal: Unified Speech-Singing Code-Switching Synthesis
Yufei Shi, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling, Yang Ai
Abstract
We propose UniVocal, a unified framework that implicitly infers vocal modes from text context to pioneer Speech-Singing Code-Switching (SCS) Synthesis - a task where transitions are autonomously driven by textual semantics, akin to seamless human language blending. Unlike single-mode generation or systems relying on switching-control tags, our proposed UniVocal implicitly infers vocal modes solely from text context. To achieve this, we employ a data-efficient two-stage curriculum learning strategy that progressively trains a competitive TTS system to acquire the desired SCS capability. Addressing data scarcity, we introduce a scalable pipeline to synthesize diverse code-switching data that is both semantically and acoustically natural, alongside a new multi-scenario benchmark, SCSBench. To address limitations of semantic tokenizers in capturing acoustic details, we also introduce refined cent token and Chain-of-Thought (CoT) generation for planning prosody before content generation, effectively enhancing empathetic speech generation and singing melody. Experimental results demonstrate that UniVocal achieves state-of-the-art performance on SCSBench while maintaining competitive performance on regular speech and singing tasks. Audio samples are available at https://project-univocal-demo.github.io/demo/. The code and dataset are released at https://github.com/FunAudioLLM/FunResearch/tree/main/UniVocal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
- UniSyn: An End-to-End Unified Model for Text-to-Speech and Singing Voice SynthesisYi Lei, Shan Yang, Xinsheng Wang, Qicong Xie et al.AAAI 2023 · 15 citations
Related papers
- UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech SynthesisXinfa Zhu, Wenjie Tian, Xinsheng Wang, Lei He et al.ACM MM 2024 · 3 citations
- UniSS: Unified Expressive Speech-to-Speech Translation with Your VoiceSitong Cheng, Bianweizhen, Xinsheng Wang, Ruibin Yuan et al.ICLR 2026 · 7 citations
- TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style ControlYu Zhang, Ziyue Jiang, Ruiqi Li, Changhao Pan et al.EMNLP 2024 · 4 citations
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang et al.ACL 2026 · 2 citations
- UniSinger: Unified End-to-End Singing Voice Synthesis With Cross-Modality Information MatchingZhiqing Hong, Chenye Cui, Rongjie Huang, Lichao Zhang et al.ACM MM 2023 · 3 citations
