ElePose: Unsupervised 3D Human Pose Estimation by Predicting Camera Elevation and Learning Normalizing Flows on 2D Poses
Bastian Wandt, James J. Little, Helge Rhodin
Abstract
Human pose estimation from single images is a challenging problem that is typically solved by supervised learning. Unfortunately, labeled training data does not yet exist for many human activities since 3D annotation requires dedicated motion capture systems. Therefore, we propose an unsupervised approach that learns to predict a 3D human pose from a single image while only being trained with 2D pose data, which can be crowd-sourced and is already widely available. To this end, we estimate the 3D pose that is most likely over random projections, with the likelihood estimated using normalizing flows on 2D poses. While previous work requires strong priors on camera rotations in the training data set, we learn the distribution of camera angles which significantly improves the performance. Another part of our contribution is to stabilize training with normalizing flows on high-dimensional 3D pose data by first projecting the 2D poses to a linear subspace. We outperform the stateof-the-art unsupervised human pose estimation methods on the benchmark datasets Human3.6M and MPI-INF-3DHP in many metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- DiffPose: Multi-hypothesis Human Pose Estimation using Diffusion ModelsKarl Holmquist, Bastian WandtICCV 2023 · 92 citations
- Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D MotionsJianan Li, Xiao Chen, Tao Huang, Tien-Tsin WongCVPR 2026 · 3 citations
- Perspose: 3D Human Pose Estimation with Perspective Encoding and Perspective RotationXiaoyang Hao, Han LiICCV 2025 · 3 citations
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D DiffusionHongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu et al.CVPR 2026 · 2 citations
- Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D PretrainingZhumei Wang, Zechen Hu, Ruoxi Guo, Huaijin Pi et al.CVPR 2026 · 1 citation
Builds on18
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- XNect: real-time multi-person 3D motion capture with a single RGB cameraDushyant Mehta, Oleksandr Sotnychenko, Franziska Mueller, Weipeng Xu et al.SIGGRAPH 2020 · 267 citations
- Probabilistic Modeling for Human Mesh RecoveryNikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, Kostas DaniilidisICCV 2021 · 201 citations
- SPEC: Seeing People in the Wild with an Estimated CameraMuhammed Kocabas, Chun-Hao P. Huang, Joachim Tesch, Lea Müller et al.ICCV 2021 · 181 citations
Related papers
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin et al.CVPR 2021
- Probabilistic Monocular 3D Human Pose Estimation with Normalizing FlowsTom Wehrbein, Marco Rudolph, Bodo Rosenhahn, Bastian WandtICCV 2021 · 147 citations
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
- Geometry-Driven Self-Supervised Method for 3D Human Pose EstimationYang Li, Kan Li, Shuai Jiang, Ziyue Zhang et al.AAAI 2020 · 40 citations
- On Boosting Single-Frame 3D Human Pose Estimation via Monocular VideosZhi Li, Xuan Wang, Fei Wang, Peilin JiangICCV 2019 · 46 citations
