Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
Zhumei Wang, Zechen Hu, Ruoxi Guo, Huaijin Pi, Ziyong Feng, Liang Zhang, Mingtao Pei, Siyuan Huang
Abstract
Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models' reliance on 3D training data from constrained environments limits their out-of-distribution generalization; and (2) the inherent difficulty of estimating metric-scale poses from monocular observations. This paper introduces Mocap-2-to-3, a novel framework that differs from prior HMR methods by recovering absolute poses from monocular input and leveraging abundant 2D data to enhance 3D motion recovery. To effectively utilize the action priors and diversity in large-scale 2D datasets, we reformulate 3D motion as a multi-view synthesis process and divide the training into two stages: a single-view diffusion model is first pre-trained on extensive 2D data, followed by multi-view fine-tuning on 3D data, thus achieving a combination of strong priors and geometric constraints. Further-more, to recover absolute poses, we introduce a novel human motion representation that decouples the learning of local pose and global movements, while encoding ground geometric priors to accelerate convergence, thereby yielding more precise positioning in the physical world. Experiments on in-the-wild benchmarks show that our method outperforms state-of-the-art approaches in both cameraspace motion realism and world-grounded human positioning, while exhibiting strong generalization capability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec04a40e-8d92-40f4-a41f-d924cc4cfce9Builds on30
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 701 citations
Related papers
- AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D DiffusionHongjie Li, Heng Yu, Jiaman Li, Hong-Xing Yu et al.CVPR 2026 · 2 citations
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin et al.CVPR 2021
- Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D MotionsJianan Li, Xiao Chen, Tao Huang, Tien-Tsin WongCVPR 2026 · 3 citations
- Neural monocular 3D human motion capture with physical awarenessSoshi Shimada, Vladislav Golyanik, Weipeng Xu, Patrick Pérez et al.SIGGRAPH 2021 · 107 citations
