H3WB: Human3.6M 3D WholeBody Dataset and Benchmark
Yue Zhu, Nermin Samet, David Picard
Abstract
We present a benchmark for 3D human whole-body pose estimation, which involves identifying accurate 3D keypoints on the entire human body, including face, hands, body, and feet. Currently, the lack of a fully annotated and accurate 3D whole-body dataset results in deep networks being trained separately on specific body parts, which are combined during inference. Or they rely on pseudo-groundtruth provided by parametric body models which are not as accurate as detection based methods. To overcome these issues, we introduce the Human3.6M 3D WholeBody (H3WB) dataset, which provides whole-body annotations for the Human3.6M dataset using the COCO Wholebody layout. H3WB comprises 133 whole-body keypoint annotations on 100K images, made possible by our new multi-view pipeline. We also propose three tasks: i) 3D whole-body pose lifting from 2D complete whole-body pose, ii) 3D whole-body pose lifting from 2D incomplete whole-body pose, and iii) 3D whole-body pose estimation from a single RGB image. Additionally, we report several baselines from popular methods for these tasks. Furthermore, we also provide automated 3D whole-body annotations of TotalCapture and experimentally show that when used with H3WB it helps to improve the performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85f38b24-1f0b-4726-af39-5949eaebb055Cited by top-tier papers6
- avaTTAR: Table Tennis Stroke Training with Embodied and Detached Visualization in Augmented RealityDizhi Ma, Xiyun Hu, Jingyu Shi, Mayank Patel et al.UIST 2024 · 21 citations
- Humoto: A 4D Dataset of Mocap Human Object InteractionsJiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang et al.ICCV 2025 · 4 citations
- Cross-View Isolated Sign Language Recognition via View Synthesis and Feature DisentanglementXin Shen, Xinyu Wang, Lei Shen, Kaihao Zhang et al.ICCV 2025 · 1 citation
- Non-rigid Structure-from-Motion: Temporally-smooth Procrustean Alignment and Spatially-variant Deformation ModelingJiawei Shi, Hui Deng, Yuchao DaiCVPR 2024 · 1 citation
- 3D-LFM: Lifting Foundation ModelMosam Dabhi, László A. Jeni, Simon LuceyCVPR 2024
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageGyeongsik Moon, Ju Yong Chang, Kyoung Mu LeeICCV 2019 · 368 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
Related papers
- MEBOW: Monocular Estimation of Body Orientation in the WildChenyan Wu, Yukun Chen, Jiajia Luo, Che-Chun Su et al.CVPR 2020
- Geometry-Driven Self-Supervised Method for 3D Human Pose EstimationYang Li, Kan Li, Shuai Jiang, Ziyue Zhang et al.AAAI 2020 · 40 citations
- Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationZhenbo Yu, Bingbing Ni, Jingwei Xu, Junjie Wang et al.ICCV 2021 · 39 citations
- UltraPose: Synthesizing Dense Pose with 1 Billion Points by Human-body Decoupling 3D ModelHaonan Yan, Jiaqi Chen, Xujie Zhang, Shengkai Zhang et al.ICCV 2021 · 16 citations
- SAM 3D Body: Robust Full-Body Human Mesh RecoveryXitong Yang, Devansh Kukreja, Don Pinkus, Taosha Fan et al.CVPR 2026 · 81 citations
