Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator
Gyeongsik Moon
Abstract
Accurately recovering hand poses within the body context remains a major challenge in 3D whole-body pose estimation. This difficulty arises from a fundamental supervision gap: whole-body pose estimators are trained on full-body datasets with limited hand diversity, while hand-only estimators, trained on hand-centric datasets, excel at detailed finger articulation but lack global body awareness. To address this, we propose WholeBody++, a modular framework that leverages the strengths of both pre-trained whole-body and hand pose estimators. We introduce CHAM (Conditional Hands Modulator), a lightweight module that modulates the whole-body feature stream using hand-specific features extracted from a pre-trained hand pose estimator. This modulation enables the whole-body model to predict wrist orientations that are both accurate and coherent with the upper-body kinematic structure, without retraining the full-body model. In parallel, we directly incorporate finger articulations and hand shapes predicted by the hand pose estimator, aligning them to the full-body mesh via differentiable rigid alignment. This design allows WholeBody++ to combine globally consistent body reasoning with fine-grained hand detail. Extensive experiments demonstrate that WholeBody++ substantially improves hand accuracy and enhances overall full-body pose quality. Code and pretrained models will be released publicly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44e4d2df-5435-44e2-8518-e4c47401fc2bBuilds on21
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- ControlVideo: Training-free Controllable Text-to-video GenerationYabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang et al.ICLR 2024 · 359 citations
Related papers
- HMR-Adapter: A Lightweight Adapter with Dual-Path Cross Augmentation for Expressive Human Mesh RecoveryWenhao Shen, Wanqi Yin, Hao Wang, Chen Wei et al.ACM MM 2024 · 2 citations
- Towards Robust and Expressive Whole-body Human Pose and Shape EstimationHui En Pang, Zhongang Cai, Lei Yang, Qingyi Tao et al.NeurIPS 2023 · 16 citations
- Expressive Forecasting of 3D Whole-Body Human MotionsPengxiang Ding, Qiongjie Cui, Haofan Wang, Min Zhang et al.AAAI 2024 · 10 citations
- Monocular Real-Time Full Body Capture With Inter-Part CorrelationsYuxiao Zhou, Marc Habermann, Ikhsanul Habibie, Ayush Tewari et al.CVPR 2021
- Forecasting of 3D Whole-Body Human Poses with Grasping ObjectsHaitao Yan, Qiongjie Cui, Jiexin Xie, Shijie GuoCVPR 2024 · 6 citations
