ProxyCap: Real-Time Monocular Full-Body Capture in World Space via Human-Centric Proxy-to-Motion Learning
Yuxiang Zhang, Hongwen Zhang, Liangxiao Hu, Jiajun Zhang, Hongwei Yi, Shengping Zhang, Yebin Liu
Abstract
Learning-based approaches to monocular motion capture have recently shown promising results by learning to regress in a data-driven manner. However, due to the challenges in data collection and network designs, it remains challenging to achieve real-time full-body capture while being accurate in world space. In this work, we introduce ProxyCap, a human-centric proxy-to-motion learning scheme to learn world-space motions from a proxy dataset of 2D skeleton sequences and 3D rotational motions. Such proxy data enables us to build a learning-based network with accurate world-space supervision while also mitigating the generalization issues. For more accurate and physically plausible predictions in world space, our network is designed to learn human motions from a human-centric perspective, which enables the understanding of the same motion captured with different camera trajectories. Moreover, a contact-aware neural motion descent module is proposed to improve foot-ground contact and motion misalignment with the proxy observations. With the proposed learning-based solution, we demonstrate the first real-time monocular full-body capture system with plausible foot-ground contact in world space even using hand-held cameras.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe7e9053-9711-4cce-8366-baa67d335dadCited by top-tier papers7
- A Plug-And-Play Physical Motion Restoration Approach for In-The-Wild High-Difficulty MotionsYouliang Zhang, Ronghui Li, Yachao Zhang, Liang Pan et al.ICCV 2025 · 4 citations
- MBTI: Masked Blending Transformers with Implicit Positional Encoding for Frame-rate Agnostic Motion EstimationJungwoo Huh, Yeseung Park, Seongjean Kim, Jungsu Kim et al.ICCV 2025
- HHMR: Holistic Hand Mesh Recovery by Enhancing the Multimodal Controllability of Graph Diffusion ModelsMengcheng Li, Hongwen Zhang, Yuxiang Zhang, Ruizhi Shao et al.CVPR 2024
- HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera DecouplingFengyuan Yang, Tanuj Sur, Tze Ho Elden Tse, Angela YaoCVPR 2026
- Lodge: A Coarse to Fine Diffusion Network for Long Dance Generation Guided by the Characteristic Dance PrimitivesRonghui Li, Yuxiang Zhang, Yachao Zhang, Hongwen Zhang et al.CVPR 2024
Builds on34
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- Learning an animatable detailed 3D face model from in-the-wild imagesYao Feng, Haiwen Feng, Michael J. Black, Timo BolkartSIGGRAPH 2021 · 662 citations
- 3D Human Pose Estimation with Spatial and Temporal TransformersCe Zheng, Sijie Zhu, Matías Mendieta, Taojiannan Yang et al.ICCV 2021 · 648 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
Related papers
- Neural monocular 3D human motion capture with physical awarenessSoshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick Pérez et al.SIGGRAPH 2021 · 107 citations
- Physics-based Human Motion Estimation and Synthesis from VideosKevin Xie, Tingwu Wang, Umar Iqbal, Yunrong Guo et al.ICCV 2021 · 102 citations
- Learning Motion Priors for 4D Human Body Capture in 3D ScenesSiwei Zhang, Yan Zhang, Federica Bogo, Marc Pollefeys et al.ICCV 2021 · 117 citations
- HybridCap: Inertia-Aid Monocular Capture of Challenging Human MotionsHan Liang, Yannan He, Chengfeng Zhao, Mutian Li et al.AAAI 2023 · 29 citations
- Monocular Real-Time Full Body Capture With Inter-Part CorrelationsYuxiao Zhou, Marc Habermann, Ikhsanul Habibie, Ayush Tewari et al.CVPR 2021
