Towards Robust and Expressive Whole-body Human Pose and Shape Estimation
Hui En Pang, Zhongang Cai, Lei Yang, Qingyi Tao, Zhonghua Wu, Tianwei Zhang, Ziwei Liu
Abstract
Whole-body pose and shape estimation aims to jointly predict different behaviors (e.g., pose, hand gesture, facial expression) of the entire human body from a monocular image. Existing methods often exhibit degraded performance under the complexity of in-the-wild scenarios. We argue that the accuracy and reliability of these models are significantly affected by the quality of the predicted bounding box, e.g., the scale and alignment of body parts. The natural discrepancy between the ideal bounding box annotations and model detection results is particularly detrimental to the performance of whole-body pose and shape estimation. In this paper, we propose a novel framework to enhance the robustness of whole-body pose and shape estimation. Our framework incorporates three new modules to address the above challenges from three perspectives: 1) Localization Module enhances the model's awareness of the subject's location and semantics within the image space. 2) Contrastive Feature Extraction Module encourages the model to be invariant to robust augmentations by incorporating contrastive loss with dedicated positive samples. 3) Pixel Alignment Module ensures the reprojected mesh from the predicted camera and body model parameters are accurate and pixel-aligned. We perform comprehensive experiments to demonstrate the effectiveness of our proposed framework on body, hands, face and whole-body benchmarks. Codebase is available at https://github.com/robosmplx/robosmplx.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de012a73-4486-4f46-b21b-3a298e13c1d3Cited by top-tier papers7
- AiOS: All-in-One-Stage Expressive Human Pose and Shape EstimationQingping Sun, Yanjun Wang, Ailing Zeng, Wanqi Yin et al.CVPR 2024 · 20 citations
- Viewpoint-Aware Visual Grounding in 3D ScenesXiangxi Shi, Zhonghua Wu, Stefan LeeCVPR 2024 · 13 citations
- Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and ModulationYinghao Wu, Chaoran Wang, Lu Yin, Shihui Guo et al.NeurIPS 2024 · 11 citations
- Fine Structure-Aware Sampling: A New Sampling Training Scheme for Pixel-Aligned Implicit Models in Single-View Human ReconstructionKennard Yanting Chan, Fayao Liu, Guosheng Lin, Chuan Sheng Foo et al.AAAI 2024 · 4 citations
- IPVTON: Image-based 3D Virtual Try-on with Image Prompt AdapterXiaojing Zhong, Zhonghua Wu, Xiaofeng Yang, Guosheng Lin et al.AAAI 2025 · 3 citations
Builds on22
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
Related papers
- Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands ModulatorGyeongsik MoonCVPR 2026 · 1 citation
- Glimpse: Geometry Learning of Multi-scale Structural Priors for 3D Pose EstimationZhenhua TANG, Jihua Peng, Yanbin Hao, Qiguang Miao et al.ICML 2026
- ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving AugmentationSiyuan Bian, Jiefeng Li, Jiasheng Tang, Cewu LuAAAI 2024 · 4 citations
- Monocular Real-Time Full Body Capture With Inter-Part CorrelationsYuxiao Zhou, Marc Habermann, Ikhsanul Habibie, Ayush Tewari et al.CVPR 2021
- Self-Supervised 3D Human Mesh Recovery from a Single Image with Uncertainty-Aware LearningGuoli Yan, Zichun Zhong, Jing HuaAAAI 2024 · 1 citation
