MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music Synchronization
Chenguang Li, Yu-Hui Wen, Liping Jing
Abstract
Realistic choreography demands simultaneous attention to rhythm and motivation. Prevailing automated dance generation methods mainly depend on musical input, overlooking the motivations that drive meaningful dance creation. Inspired by the motivation choreography, we aim to articulate dance motivations through textual guidance. However, the absence of high-quality datasets concurrently containing music, textual descriptions, and motion data presents a challenge in achieving accurate fine-grained textual control. To address this limitation, we present MotivDance, a novel framework integrating fine-grained textual guidance with music to synthesize semantically coherent dance sequences. Our approach first synthesizes text-guided key poses as motivations. We then introduce an Adaptive Keyframe Locator that dynamically positions these motivations within the musical context through beat-aware synchronization and cross-modal latent space alignment. Finally, a Transformer-based U-Net diffusion model performs the motion in-betweening while preserving motivational integrity. Extensive qualitative and quantitative experiments demonstrate that MotivDance effectively integrates music with fine-grained text control to generate high-fidelity dance motions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- Self-Supervised MultiModal Versatile NetworksJean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelovic et al.NeurIPS 2020 · 423 citations
- Flexible Diffusion Modeling of Long VideosWilliam Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach et al.NeurIPS 2022 · 384 citations
- Guided Motion Diffusion for Controllable Human Motion SynthesisKorrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu TangICCV 2023 · 240 citations
Related papers
- M2PE-Diff: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion FrameworkNokap Tony ParkACM MM 2025 · 2 citations
- Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion PriorForam Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong et al.AAAI 2026 · 4 citations
- X-Dancer: Expressive Music to Human Dance Video GenerationZeyuan Chen, Hongyi Xu, Guoxian Song, You Xie et al.ICCV 2025 · 7 citations
- MusicInfuser: Making Video Diffusion Listen and DanceSusung Hong, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. SeitzCVPR 2026 · 5 citations
- MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and CorrespondenceFuming You, Minghui Fang, Li Tang, Rongjie Huang et al.NeurIPS 2024 · 8 citations
