UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-Experts
Zhenyu Liu, Yunxin Li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He
Abstract
Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. However, the auditory domain remains a significant challenge, with music and speech often developed in isolation, hindering progress towards universal audio synthesis. This separation stems from inherent task conflicts between semantic speech and structural music modeling, and severe data imbalances, which impede the development of a truly unified model. To address these challenges, we propose UniMoE-Audio, a unified speech and music generation model built upon a novel Dynamic-Capacity Mix-of-Experts (DCMoE) framework. Architecturally, UniMoE-Audio extends the conventional MoE paradigm by introducing a Top-P routing strategy for adaptive capacity allocation. To tackle data imbalance, we introduce a three-stage training curriculum: 1) Independent Specialist Training leverages original datasets to instill domain-specific knowledge into each specialists without interference; 2) MoE Integration and Warmup incorporates these specialists into the UniMoE-Audio architecture, warming up the gate module and shared expert using a subset of balanced dataset; and 3) Synergistic Joint Training trains the entire model end-to-end on the fully balanced dataset, fostering enhanced cross-domain synergy. Extensive experiments show that UniMoE-Audio not only achieves state-of-the-art performance on major speech and music generation benchmarks, but also demonstrates superior synergistic learning, mitigating the performance degradation typically seen in naive joint training. Our findings highlight the substantial potential of specialized MoE architecture and curated training strategies in advancing universal audio generation. The source code is available at https: //github.com/HITsz-TMG/Uni-MoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar et al.NeurIPS 2023 · 910 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- YuE: Scaling Open Foundation Models for Long-Form Music GenerationRuibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang et al.ICLR 2026 · 112 citations
Related papers
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang et al.SIGGRAPH 2026
- UniMuMo: Unified Text, Music, and Motion GenerationHan Yang, Kun Su, Yutong Zhang, Jiaben Chen et al.AAAI 2025 · 3 citations
- MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality AssessmentHuangbiao Xu, Huanqi Wu, Xiao Ke, Junyi Wu et al.AAAI 2026 · 2 citations
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang et al.ACL 2026 · 2 citations
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio GenerationYan Rong, Jinting Wang, Guangzhi Lei, Shan Yang et al.ACM MM 2025 · 1 citation
