TravelNet: Self-supervised Physically Plausible Hand Motion Learning from Monocular Color Images
Zimeng Zhao, Xi Zhao, Yangang Wang
摘要
This paper aims to reconstruct physically plausible hand motion from monocular color images. Existing frame-by-frame estimating approaches can not guarantee the physical plausibility (e.g. penetration, jittering) directly. In this paper, we embed physical constraints on the per-frame estimated motions in both spatial and temporal space. Our key idea is to adopt a self-supervised learning strategy to train a novel encoder-decoder, named TravelNet, whose training motion data is prepared by the physics engine using discrete pose states. TravelNet captures key pose states from hand motion sequences as compact motion descriptors, inspired by the concept of keyframes in animation. Finally, it manages to extract those key states out of perturbations without manual annotations, and reconstruct the motions preserving details and physical plausibility. In the experiments, we show that the outputs of the TravelNet contain both finger synergism and time consistency. Through the proposed framework, hand motions can be accurately reconstructed and flexibly re-edited, which is superior to the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- MobRecon: Mobile-Friendly Hand Mesh Reconstruction from Monocular ImageXingyu Chen, Yufeng Liu, Yajiao Dong, Xiong Zhang 等CVPR 2022 · 被引用 97 次
- Stability-driven Contact Reconstruction From Monocular Color ImagesZimeng Zhao, Binghui Zuo, Wei Xie, Yangang WangCVPR 2022 · 被引用 15 次
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu 等CVPR 2024 · 被引用 12 次
- SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View AdaptationYinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing FuAAAI 2024 · 被引用 5 次
- Semi-Supervised Hand Appearance Recovery via Structure Disentanglement and Dual Adversarial DiscriminationZimeng Zhao, Binghui Zuo, Zhiyu Long, Yangang WangCVPR 2023
它引用的顶会 Paper7
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell 等ICCV 2019 · 被引用 493 次
- Robust motion in-betweeningFélix G. Harvey, Mike Yurick, Derek Nowrouzezahrai, Christopher J. PalSIGGRAPH 2020 · 被引用 269 次
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang 等ICCV 2019 · 被引用 248 次
- RigNet: neural rigging for articulated charactersZhan Xu, Yang Zhou, Evangelos Kalogerakis, Chris Landreth 等SIGGRAPH 2020 · 被引用 127 次
- Aligning Latent Spaces for 3D Hand Pose EstimationLinlin Yang, Shile Li, Dongheui Lee, Angela YaoICCV 2019 · 被引用 93 次
相关 Paper
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 被引用 14 次
- Physics-based Human Motion Estimation and Synthesis from VideosKevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo 等ICCV 2021 · 被引用 102 次
- Diffusion-Based 3D Hand Motion Recovery with Intuitive PhysicsYufei Zhang, Zijun Cui, Jeffrey O. Kephart, Qiang JiICCV 2025 · 被引用 1 次
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
