UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
Chunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang
Abstract
Generative audio modeling has largely been fragmented into specialized tasks, text-tospeech (TTS), text-to-music (TTM), and textto-audio (TTA), each operating under heterogeneous control paradigms. Unifying these modalities remains a fundamental challenge due to the intrinsic dissonance between structured semantic representations (speech/music) and unstructured acoustic textures (sound effects). In this paper, we introduce UniSonate, a unified flow-matching framework capable of synthesizing speech, music, and sound effects through a standardized, referencefree natural language instruction interface. To reconcile structural disparities, we propose a novel dynamic token injection mechanism that projects unstructured environmental sounds into a structured temporal latent space, enabling precise duration control within a phoneme-driven Multimodal Diffusion Transformer (MM-DiT). Coupled with a multi-stage curriculum learning strategy, this approach effectively mitigates cross-modal optimization conflicts. Extensive experiments demonstrate that UniSonate achieves state-of-the-art performance in instruction-based TTS (WER 1.47%) and TTM (SongEval Coherence 3.18), while maintaining competitive fidelity in TTA. Crucially, we observe positive transfer, where joint training on diverse audio data significantly enhances structural coherence and prosodic expressiveness compared to single-task baselines. Audio samples are available at https: //qiangchunyu.github.io/UniSonate/ . * Corresponding author. The name "Sonate" is derived from the musical term "Sonata", symbolizing the model's comprehensive capabilities in audio generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b66a21f-8be5-411b-9caa-47d995efdaecBuilds on9
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang et al.ICLR 2026 · 112 citations
- Text-to-Audio Generation using Instruction Guided Latent Diffusion ModelDeepanway Ghosal, Navonil Majumder, Ambuj Mehrish, Soujanya PoriaACM MM 2023 · 90 citations
- IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-SpeechSiyi Zhou, Yiquan Zhou, Yi He, Xun Zhou et al.AAAI 2026 · 63 citations
Related papers
- ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation AlignmentJun-Hak Yun, Seung-Bin Kim, Seong-Whan LeeACL 2026
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint TrainingTianqiao Liu, Xueyi Li, Hao Wang, Haoxuan Li et al.ICLR 2026 · 6 citations
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang et al.SIGGRAPH 2026
- OmniSonic: Towards Universal and Holistic Audio Generation from Video and TextWeiguo Pian, Saksham Singh Kushwaha, Zhimin Chen, Shijian Deng et al.CVPR 2026 · 2 citations
