DanceMVP: Self-Supervised Learning for Multi-Task Primitive-Based Dance Performance Assessment via Transformer Text Prompting
Yun Zhong, Yiannis Demiris
Abstract
Dance is generally considered to be complex for most people as it requires coordination of numerous body motions and accurate responses to the musical content and rhythm. Studies on automatic dance performance assessment could help people improve their sensorimotor skills and promote research in many fields, including human motion analysis and motion generation. Recent papers on dance performance assessment usually evaluate simple dance motions with a single task -estimating final performance scores. In this paper, we propose DanceMVP: multi-task dance performance assessment via text prompting that solves three related tasks -(i) dance vocabulary recognition, (ii) dance performance scoring and (iii) dance rhythm evaluation. In the pre-training phase, we contrastively learn the primitive-based features of complex dance motion and music using the InfoNCE loss. For the downstream task, we propose a transformer-based text prompter to perform multi-task evaluations for the three proposed assessment tasks. Also, we build a multimodal dance-music dataset named ImperialDance. The novelty of our ImperialDance is that it contains dance motions for diverse expertise levels and a significant amount of repeating dance sequences for the same choreography to keep track of the dance performance progression. Qualitative results show that our pre-trained feature representation could cluster dance pieces for different dance genres, choreographies, expertise levels and primitives, which generalizes well on both ours and other dance-music datasets. The downstream experiments demonstrate the robustness and improvement of our method over several ablations and baselines across all three tasks, as well as monitoring the users' dance level progression.
- Yiannis Demiris is supported by a Royal Academy of Engineering Chair in Emerging Technologies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion TransformerBuyu Li, Yongchi Zhao, Zhelun Shi, Lu ShengAAAI 2022 · 182 citations
Related papers
- DanceEditor: Towards Iterative Editable Music-Driven Dance Generation with Open-Vocabulary DescriptionsHengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li et al.ICCV 2025 · 3 citations
- DanceCamera3D: 3D Camera Movement Synthesis with Music and DanceZixuan Wang, Jia Jia, Shikun Sun, Haozhe Wu et al.CVPR 2024 · 5 citations
- TM2D: Bimodality Driven 3D Dance Generation via Music-Text IntegrationKehong Gong, Dongze Lian, Heng Chang, Chuan Guo et al.ICCV 2023 · 103 citations
- MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPrerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta et al.ICCV 2025 · 1 citation
- Self-supervised Dance Video Synthesis Conditioned on MusicXuanchi Ren, Haoran Li, Zijian Huang, Qifeng ChenACM MM 2020 · 68 citations
