Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition *
Xiaodan Hu, Narendra Ahuja
Abstract
Dance experts often view dance as a hierarchy of information, spanning low-level (raw images, image sequences), mid-levels (human poses and bodypart movements), and high-level (dance genre). We propose a Hierarchical Dance Video Recognition framework (HDVR). HDVR estimates 2D pose sequences, tracks dancers, and then simultaneously estimates corresponding 3D poses and 3D-to-2D imaging parameters, without requiring ground truth for 3D poses. Unlike most methods that work on a single person, our tracking works on multiple dancers, under occlusions. From the estimated 3D pose sequence, HDVR extracts body part movements, and therefrom dance genre. The resulting hierarchical dance representation is explainable to experts. To overcome noise and interframe correspondence ambiguities, we enforce spatial and temporal motion smoothness and photometric continuity over time. We use an LSTM network to extract 3D movement subsequences from which we recognize dance genre. For experiments, we have identified 154 movement types, of 16 body parts, and assembled a new University of Illinois Dance (UID) Dataset, containing 1143 video clips of 9 genres covering 30 hours, annotated with movement and genre labels. Our experimental results demonstrate that our algorithms outperform the state-of-the-art 3D pose estimation methods, which also enhances our dance recognition performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74eb7dd7-c325-47f4-99b1-43d7ba5eca4aCited by top-tier papers3
- PoseTriplet: Co-evolving 3D Human Pose Estimation, Imitation, and Hallucination under Self-supervisionKehong Gong, Bingbing Li, Jianfeng Zhang, Tao Wang et al.CVPR 2022 · 40 citations
- EgoMusic-Driven Human Dance Motion Estimation with Skeleton MambaQuang Nguyen, Nhat Le, Baoru Huang, Minh Nhat Vu et al.ICCV 2025 · 4 citations
- Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose EstimationZiyu Wang, Shuangpeng Han, Mengmi ZhangICLR 2026 · 3 citations
Builds on4
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- VIBE: Video Inference for Human Body Pose and Shape EstimationMuhammed Kocabas, Nikos Athanasiou, Michael J. BlackCVPR 2020
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi et al.CVPR 2020
- Deep Kinematics Analysis for Monocular 3D Human Pose EstimationJingwei Xu, Zhenbo Yu, Bingbing Ni, Jiancheng Yang et al.CVPR 2020
Related papers
- Disentangled Concepts Speak Louder Than Words: Explainable Video Action RecognitionJongseo Lee, Wooil Lee, Gyeong-Moon Park, Seong Tae Kim et al.NeurIPS 2025 · 4 citations
- FineDance: A Fine-grained Choreography Dataset for 3D Full Body Dance GenerationRonghui Li, Junfan Zhao, Yachao Zhang, Mingyang Su et al.ICCV 2023 · 110 citations
- DanceFormer: Music Conditioned 3D Dance Generation with Parametric Motion TransformerBuyu Li, Yongchi Zhao, Zhelun Shi, Lu ShengAAAI 2022 · 182 citations
- DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human PoseHuangbiao Xu, Xiao Ke, Huanqi Wu, Rui Xu et al.AAAI 2025 · 8 citations
- MultiPly: Reconstruction of Multiple People from Monocular Video in the WildZeren Jiang, Chen Guo, Manuel Kaufmann, Tianjian Jiang et al.CVPR 2024
