PPR: Physically Plausible Reconstruction from Monocular Videos
Gengshan Yang, Shuo Yang, John Z. Zhang, Zachary Manchester, Deva Ramanan
Abstract
Given casually-captured monocular videos (left), PPR builds 3D models of articulated objects and the surrounding environment. Naive kinematic reconstruction (middle) generates a family of solutions, some containing inconsistent physical support and contact dynamics (blue and green color), such as floating or walking with sliding feet. We show that differentiable physics simulation acts as effective regularizer for improving the physical plausibility of visual reconstruction algorithms. As PPR reconstructs the dynamics scene, it also drives a ragdoll in a physics simulator to track the kinematic reconstruction. This ensures the reconstructions are statically stable with ground contact (right), and the center of mass is projected within the support polygon (marked with red). PPR also reports physics estimations, such as ground reaction forces (red arrows) and center of mass (green arrow).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce6adba3-c661-430b-97d8-cda8da7863fbCited by top-tier papers18
- PhysPT: Physics-aware Pretrained Transformer for Estimating Human Dynamics from Monocular VideosYufei Zhang, Jeffrey O. Kephart, Zijun Cui, Qiang JiCVPR 2024 · 14 citations
- Seeing the Wind from a Falling LeafZhiyuan Gao, Jiageng Mao, Hong-Xing Yu, Haozhe Lou et al.NeurIPS 2025 · 10 citations
- VAREN: Very Accurate and Realistic Equine NetworkSilvia Zuffi, Ylva Mellbin, Ci Li, Markus Höschle et al.CVPR 2024 · 9 citations
- HoliGS: Holistic Gaussian Splatting for Embodied View SynthesisXiaoyuan Wang, Yizhou Zhao, Botao Ye, Xiaojun Shan et al.NeurIPS 2025 · 8 citations
- REACTO: Reconstructing Articulated Objects from a Single VideoChaoyue Song, Jiacheng Wei, Chuan Sheng Foo, Guosheng Lin et al.CVPR 2024 · 7 citations
Builds on35
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 1,421 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
Related papers
- Differentiable Dynamics for Articulated 3d Human Motion ReconstructionErik Gärtner, Mykhaylo Andriluka, Erwin Coumans, Cristian SminchisescuCVPR 2022 · 33 citations
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- SAFT: Shape and Appearance of Fabrics from Template via Differentiable Physical Simulations from Monocular VideoDavid Stotko, Reinhard KleinICCV 2025 · 1 citation
- Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular VideoErik Gärtner, Mykhaylo Andriluka, Hongyi Xu, Cristian SminchisescuCVPR 2022 · 31 citations
- gradSim: Differentiable simulation for system identification and visuomotor controlJ. Krishna Murthy, Miles Macklin, Florian Golemo, Vikram Voleti et al.ICLR 2021 · 130 citations
