Body2Hands: Learning To Infer 3D Hands From Conversational Gesture Body Dynamics
Evonne Ng, Shiry Ginosar, Trevor Darrell, Hanbyul Joo
Abstract
We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in non-verbal communication settings. We formulate the learning of this prior as a prediction task of 3D hand shape over time given body motion input alone. Trained with 3D pose estimations obtained from a large-scale dataset of internet videos, our hand prediction model produces convincing 3D hand gestures given only the 3D motion of the speaker’s arms as input. We demonstrate the efficacy of our method on hand gesture synthesis from body motion input, and as a strong body prior for single-view image-based 3D hand pose estimation. We demonstrate that our method outperforms previous state-of-the-art approaches and can generalize beyond the monologue-based training data to multiperson conversations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- MobRecon: Mobile-Friendly Hand Mesh Reconstruction from Monocular ImageXingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang et al.CVPR 2022 · 97 citations
- Learning to Listen: Modeling Non-Deterministic Dyadic Facial MotionEvonne Ng, Hanbyul Joo, Liwen Hu, Hao Li et al.CVPR 2022 · 87 citations
- Driving-signal aware full-body avatarsTimur M. Bagautdinov, Chenglei Wu, Tomas Simon, Fabián Prada et al.SIGGRAPH 2021 · 71 citations
- BEST: BERT Pre-training for Sign Language Recognition with Coupling TokenizationWeichao Zhao, Hezhen Hu, Wengang Zhou, Jiaxin Shi et al.AAAI 2023 · 70 citations
- EMAGE: Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture ModelingHaiyang Liu, Zihao Zhu, Giorgio Becherini, Yichen Peng et al.CVPR 2024 · 55 citations
Builds on7
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and SynthesisGilwoo Lee, Zhiwei Deng, Shugao Ma, Takaaki Shiratori et al.ICCV 2019 · 114 citations
- Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal DataYuxiao Zhou, Marc Habermann, Weipeng Xu, Ikhsanul Habibie et al.CVPR 2020
Related papers
- Spatial-Temporal Parallel Transformer for Arm-Hand Dynamic EstimationShuying Liu, Wenbin Wu, Jiaxian Wu, Yue LinCVPR 2022 · 9 citations
- Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional NetworksYujun Cai, Liuhao Ge, Jun Liu, Jianfei Cai et al.ICCV 2019 · 504 citations
- Reconstructing Signing Avatars from Video Using Linguistic PriorsMaria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas et al.CVPR 2023
- Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose EstimationPengfei Ren, Jingyu Wang, Haifeng Sun, Qi Qi et al.ICCV 2025 · 1 citation
- Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional ControlZunnan Xu, Yachao Zhang, Sicheng Yang, Ronghui Li et al.AAAI 2024 · 20 citations
