RapVerse: Coherent Vocals and Whole-Body Motion Generation from Text
Jiaben Chen, Xin Yan, Yihang Chen, Siyuan Cen, Zixin Wang, Qinwei Ma, Haoyu Zhen, Kaizhi Qian, Lie Lu, Chuang Gan
Abstract
In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dataset, a large dataset containing synchronous rapping vocals, lyrics, and high-quality 3D holistic body meshes. With the RapVerse dataset, we investigate the extent to which scaling autoregressive multimodal transformers across language, audio, and motion can enhance the coherent and realistic generation of vocals and whole-body human motions. For modality unification, a vector-quantized variational autoencoder is employed to encode whole-body motion sequences into discrete motion tokens, while a vocal-to-unit model is leveraged to obtain quantized audio tokens preserving content, prosodic information and singer identity. By jointly performing transformer modeling on these three modalities in a unified way, our framework ensures a seamless and realistic blend of vocals and human motions. Extensive experiments demonstrate that our unified generation framework not only produces coherent and realistic singing vocals alongside human motions directly from textual inputs, but also rivals the performance of specialized single-modality generation systems, establishing new benchmarks for joint vocal-motion generation. Video demonstration and more information can be found on the project page11https://jiabenchen.github.io/RapVerse/..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b6617d4-185a-4bfc-a6cd-23c3e442be62Cited by top-tier papers4
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyJuze Zhang, Changan Chen, Xin Chen, Heng Yu et al.CVPR 2026 · 7 citations
- DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion ModelsQichao Wang, Yunhong Lu, Hengyuan Cao, Junyi Zhang et al.CVPR 2026 · 4 citations
- MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual TripletsJinfeng Xu, Zheyu Chen, Jinze Li, Shuo Yang et al.KDD 2025 · 4 citations
- Spherical Geometry Diffusion: Generating High-quality 3D Face Geometry via Sphere-anchored RepresentationsJunyi Zhang, Yiming Wang, Yunhong Lu, Qichao Wang et al.AAAI 2026 · 1 citation
Builds on22
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin et al.ICLR 2021 · 513 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
Related papers
- UniMuMo: Unified Text, Music, and Motion GenerationHan Yang, Kun Su, Yutong Zhang, Jiaben Chen et al.AAAI 2025 · 3 citations
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationKehong Gong, Dongze Lian, Heng Chang, Chuan Guo et al.ICCV 2023 · 103 citations
- DeepRapper: Neural Rap Generation with Rhyme and Rhythm ModelingLanqing Xue, Kaitao Song, Duocai Wu, Xu Tan et al.ACL 2021
- Drop the Beat! Freestyler for Accompaniment Conditioned Rapping Voice GenerationZiqian Ning, Shuai Wang, Yuepeng Jiang, Jixun Yao et al.AAAI 2025 · 5 citations
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang et al.ICLR 2026 · 56 citations
