M2PE-Diff: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion Framework
Nokap Tony Park
Abstract
Automated choreography generation, which aims to seamlessly harmonize human movements with music, is a multifaceted challenge demanding both technical precision and artistic expressiveness. We present M2PE-DIFF, a novel framework for generating human dance videos conditioned on a reference image and music sequence using a latent diffusion model. Our approach integrates a Music-to-Pose Encoder (M2PEnc), trained with a novel synthetic dataset generation pipeline (SDGPip), which maps audio features into structured 3D pose and shape parameters that capture human geometry and dynamic motion patterns synchronized with musical input. By combining these encoded parameters with a reference image through a multi-level attention mechanism within the latent diffusion framework, we synthesize visually coherent and rhythmically synchronized dance animations of individuals depicted in the given reference image. Experiments on benchmark datasets demonstrate that M2PE-DIFF achieves state-of-the-art performance, producing high-quality dance videos that accurately reflect pose diversity and temporal consistency. Additionally, our method exhibits robust generalization capabilities, validated by its strong performance on a newly introduced in-the-wild dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 711fb0b6-b637-412b-93bf-ad81364a09d1Related papers
- X-Dancer: Expressive Music to Human Dance Video GenerationZeyuan Chen, Hongyi Xu, Guoxian Song, You Xie et al.ICCV 2025 · 7 citations
- MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music SynchronizationChenguang Li, Yu-Hui Wen, Liping JingAAAI 2026
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video GenerationKaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng et al.SIGGRAPH 2026 · 3 citations
- ChoreoMuse: Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent MotionXuanchen Wang, Heng Wang, Weidong CaiACM MM 2025 · 5 citations
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D PosesYatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng et al.ICCV 2025 · 2 citations
