AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
Ruilong Li, Shan Yang, David A. Ross, Angjoo Kanazawa
Abstract
We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 5.2 hours of 3D dance motion in 1408 sequences, covering 10 dance genres with multi-view videos with known camera poses—the largest dataset of this kind to our knowledge. We show that naively applying sequence models such as transformers to this dataset for the task of music conditioned 3D motion generation does not produce satisfactory 3D motion that is well correlated with the input music. We overcome these shortcomings by introducing key changes in its architecture design and supervision: FACT model involves a deep cross-modal transformer block with full-attention that is trained to predict N future motions. We empirically show that these changes are key factors in generating long sequences of realistic dance motion that are well-attuned to the input music. We conduct extensive experiments on AIST++ with user studies, where our method outperforms recent state-of-the-art methods both qualitatively and quantitatively. The code and the dataset can be found at: https://google.github.io/aichoreographer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab7ac305-91fc-4190-80a4-5adc72144c8cCited by top-tier papers221
- Video Action DifferencingJames Burgess, Xiaohan Wang, Yuhui Zhang, Anita Rau et al.ICLR 2025 · 1,149 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- Guided Motion Diffusion for Controllable Human Motion SynthesisKorrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu TangICCV 2023 · 240 citations
- OmniControl: Control Any Joint at Any Time for Human Motion GenerationYiming Xie, Varun Jampani, Lei Zhong, Deqing Sun et al.ICLR 2024 · 228 citations
Builds on14
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Entangled Transformer for Image CaptioningGuang Li, Linchao Zhu, Ping Liu, Yi YangICCV 2019 · 346 citations
- Cross View Fusion for 3D Human Pose EstimationHaibo Qiu, Chunyu Wang, Jingdong Wang, Naiyan Wang et al.ICCV 2019 · 242 citations
- Human Motion Prediction via Spatio-Temporal InpaintingAlejandro Hernandez Ruiz, Jürgen Gall, Francesc MorenoICCV 2019 · 233 citations
Related papers
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion TransformerBuyu Li, Yongchi Zhao, Zhelun Shi, Lu ShengAAAI 2022 · 182 citations
- DanceCamera3D: 3D Camera Movement Synthesis with Music and DanceZixuan Wang, Jia Jia, Shikun Sun, Haozhe Wu et al.CVPR 2024 · 5 citations
- MEGADance: Mixture-of-Experts Architecture for Genre-Aware 3D Dance GenerationKaixing Yang, Xulong Tang, Ziqiao Peng, Yuxuan Hu et al.NeurIPS 2025 · 25 citations
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationRonghui Li, Junfan Zhao, Yachao Zhang, Mingyang Su et al.ICCV 2023 · 110 citations
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationKehong Gong, Dongze Lian, Heng Chang, Chuan Guo et al.ICCV 2023 · 103 citations
