UniMuMo: Unified Text, Music, and Motion Generation
Han Yang, Kun Su, Yutong Zhang, Jiaben Chen, Kaizhi Qian, Gaowen Liu, Chuang Gan
Abstract
We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired music and motion data based on rhythmic patterns to leverage existing large-scale music-only and motion-only datasets. By converting music, motion, and text into token-based representation, our model bridges these modalities through a unified encoder-decoder transformer architecture. To support multiple generation tasks within a single framework, we introduce several architectural improvements. We propose encoding motion with a music codebook, mapping motion into the same feature space as music. We introduce a music-motion parallel generation scheme that unifies all music and motion generation tasks into a single transformer decoder architecture with a single training task of music-motion joint generation. Moreover, the model is designed by fine-tuning existing pre-trained single-modality models, significantly reducing computational demands. Extensive experiments demonstrate that UniMuMo achieves competitive results on all unidirectional generation benchmarks across music, motion, and text modalities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c08931e-22a3-4ccb-bf95-74fc3cdf9d71Cited by top-tier papers7
- 🎧MOSPA: Human Motion Generation Driven by Spatial AudioShuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan et al.NeurIPS 2025 · 13 citations
- MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion ParadigmZiyan Guo, Zeyu Hu, De Wen Soh, Na ZhaoICCV 2025 · 10 citations
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyJuze Zhang, Changan Chen, Xin Chen, Heng Yu et al.CVPR 2026 · 7 citations
- Next-Scale Autoregressive Models for Text-to-Motion GenerationZhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin ZhaoCVPR 2026 · 6 citations
- DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsHengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li et al.ICCV 2025 · 3 citations
Builds on12
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
Related papers
- MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and CorrespondenceFuming You, Minghui Fang, Li Tang, Rongjie Huang et al.NeurIPS 2024 · 8 citations
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationKehong Gong, Dongze Lian, Heng Chang, Chuan Guo et al.ICCV 2023 · 103 citations
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive LearningWei Li, Can Gao, Guocheng Niu, Xinyan Xiao et al.ACL 2021
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang et al.CVPR 2026 · 10 citations
- RapVerse: Coherent Vocals and Whole-Body Motion Generation from TextJiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen et al.ICCV 2025 · 7 citations
