DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
Qiaosong Qi, Le Zhuo, Aixi Zhang, Yue Liao, Fei Fang, Si Liu, Shuicheng Yan
摘要
When hearing music, it is natural for people to dance to its rhythm. Automatic dance generation, however, is a challenging task due to the physical constraints of human motion and rhythmic alignment with target music. Conventional autoregressive methods introduce compounding errors during sampling and struggle to capture the long-term structure of dance sequences. To address these limitations, we present a novel cascaded motion diffusion model, DiffDance, designed for high-resolution, long-form dance generation. This model comprises a music-to-dance diffusion model and a sequence super-resolution diffusion model. To bridge the gap between music and motion for conditional generation, DiffDance employs a pretrained audio representation learning model to extract music embeddings and further align its embedding space to motion via contrastive loss. During training our cascaded diffusion model, we also incorporate multiple geometric losses to constrain the model outputs to be physically plausible and add a dynamic loss weight that adaptively changes over diffusion timesteps to facilitate sample diversity. Through comprehensive experiments performed on the benchmark dataset AIST++, we demonstrate that DiffDance is capable of generating realistic dance sequences that align effectively with the input music. These results are comparable to those achieved by state-of-the-art autoregressive methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- VideoVLA: Video Generators Can Be Generalizable Robot ManipulatorsYichao Shen, Fangyun Wei, Zhiying Du, Yaobo Liang 等NeurIPS 2025 · 被引用 73 次
- SoPo: Text-to-Motion Generation Using Semi-Online Preference OptimizationXiaofeng Tan, Hongsong Wang, Xin Geng, Pan ZhouNeurIPS 2025 · 被引用 18 次
- MusicInfuser: Making Video Diffusion Listen and DanceSusung Hong, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. SeitzCVPR 2026 · 被引用 5 次
- ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent MotionXuanchen Wang, Heng Wang, Weidong CaiACM MM 2025 · 被引用 5 次
- Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationCongyi Fan, Jian Guan, Xuanjia Zhao, Dongli Xu 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- Bidirectional Autoregressive Diffusion Model for Dance GenerationCanyu Zhang, Youbao Tang, Ning Zhang, Ruei-Sung Lin 等CVPR 2024 · 被引用 9 次
- Dance Revolution: Long-Term Dance Generation with Music via Curriculum LearningRuozi Huang, Huang Hu, Wei Wu, Kei Sawada 等ICLR 2021 · 被引用 147 次
- M2PE-Diff: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion FrameworkNokap Tony ParkACM MM 2025 · 被引用 2 次
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationRonghui Li, Junfan Zhao, Yachao Zhang, Mingyang Su 等ICCV 2023 · 被引用 110 次
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video GenerationKaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 等SIGGRAPH 2026 · 被引用 3 次
