TravelNet: Self-supervised Physically Plausible Hand Motion Learning from Monocular Color Images
Zimeng Zhao, Xi Zhao, Yangang Wang
Abstract
This paper aims to reconstruct physically plausible hand motion from monocular color images. Existing frame-by-frame estimating approaches can not guarantee the physical plausibility (e.g. penetration, jittering) directly. In this paper, we embed physical constraints on the per-frame estimated motions in both spatial and temporal space. Our key idea is to adopt a self-supervised learning strategy to train a novel encoder-decoder, named TravelNet, whose training motion data is prepared by the physics engine using discrete pose states. TravelNet captures key pose states from hand motion sequences as compact motion descriptors, inspired by the concept of keyframes in animation. Finally, it manages to extract those key states out of perturbations without manual annotations, and reconstruct the motions preserving details and physical plausibility. In the experiments, we show that the outputs of the TravelNet contain both finger synergism and time consistency. Through the proposed framework, hand motions can be accurately reconstructed and flexibly re-edited, which is superior to the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- MobRecon: Mobile-Friendly Hand Mesh Reconstruction from Monocular ImageXingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang et al.CVPR 2022 · 97 citations
- Stability-driven Contact Reconstruction From Monocular Color ImagesZimeng Zhao, Binghui Zuo, Wei Xie, Yangang WangCVPR 2022 · 15 citations
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu et al.CVPR 2024 · 12 citations
- SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationYinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing FuAAAI 2024 · 5 citations
- Semi-Supervised Hand Appearance Recovery via Structure Disentanglement and Dual Adversarial DiscriminationZimeng Zhao, Binghui Zuo, Zhiyu Long, Yangang WangCVPR 2023
Builds on7
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 269 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- RigNet: neural rigging for articulated charactersZhan Xu, Yang Zhou, Evangelos Kalogerakis, Chris Landreth et al.SIGGRAPH 2020 · 127 citations
- Aligning Latent Spaces for 3D Hand Pose EstimationLinlin Yang, Shile Li, Dongheui Lee, Angela YaoICCV 2019 · 93 citations
Related papers
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- Physics-based Human Motion Estimation and Synthesis from VideosKevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo et al.ICCV 2021 · 102 citations
- Diffusion-Based 3D Hand Motion Recovery with Intuitive PhysicsYufei Zhang, Zijun Cui, Jeffrey O. Kephart, Qiang JiICCV 2025 · 1 citation
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
