X-Dancer: Expressive Music to Human Dance Video Generation
Zeyuan Chen, Hongyi Xu, Guoxian Song, You Xie, Chenxu Zhang, Xin Chen, Chao Wang, Di Chang, Linjie Luo
Abstract
We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-diffusion framework, featuring an autoregressive transformer model that synthesize extended and music-synchronized token sequences for 2D body, head and hands poses, which then guide a diffusion model to produce coherent and realistic dance video frames. Unlike traditional methods that primarily generate human motion in 3D, X-Dancer addresses data limitations and enhances scalability by modeling a wide spectrum of 2D dance motions, capturing their nuanced alignment with musical beats through readily available monocular videos. To achieve this, we first build a spatially compositional token representation from 2D human pose labels associated with keypoint confidences, encoding both large articulated body movements (e.g., upper and lower body) and fine-grained motions (e.g., head and hands). We then design a music-to-motion transformer model that autoregressively generates music-aligned dance pose token sequences, incorporating global attention to both musical style and prior motion context. Finally we leverage a diffusion backbone to animate the reference image with these synthesized pose tokens through AdaIN, forming a fully differentiable end-to-end framework. Experimental results demonstrate that X-Dancer is able to produce both diverse and characterized dance videos, substantially outperforming state-of-the-art methods in term of diversity, expressiveness and realism. Code and model will be available for research purposes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ded631a2-3785-4888-a066-2e8ea941787cCited by top-tier papers4
- Audio-Sync Video Generation with Multi-Stream Temporal ControlShuchen Weng, Haojie Zheng, Zheng Chang, Si Li et al.NeurIPS 2025 · 14 citations
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video GenerationKaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng et al.SIGGRAPH 2026 · 3 citations
- TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-GenerationXiaoda Yang, Majun Zhang, Changhao Pan, Nick Huang et al.ICML 2026 · 1 citation
- Animate and Sound an ImageXihua Wang, Ruihua Song, Chongxuan Li, Xin Cheng et al.CVPR 2025
Builds on32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- M2PE-Diff: Music-to-Pose Encoder for Dance Video Generation Leveraging Latent Diffusion FrameworkNokap Tony ParkACM MM 2025 · 2 citations
- MotivDance: Fine-Grained Text-Guided Motivation Choreography with Music SynchronizationChenguang Li, Yu-Hui Wen, Liping JingAAAI 2026
- Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet SupervisionHyunsoo Cha, Wonjung Woo, Byungjun Kim, Hanbyul JooCVPR 2026 · 1 citation
- Controllable 3D Dance Generation Using Diffusion-Based Transformer U-NetPuyuan Guo, Tuo Hao, Wenxin Fu, Yingming Gao et al.AAAI 2025 · 5 citations
- MultiAnimate: Pose-Guided Image Animation Made ExtensibleYingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An et al.CVPR 2026 · 6 citations
