Recovering Physically Plausible Human-Object Interactions from Monocular Videos
Dingbang Huang, Etienne Vouga, Qixing Huang, Georgios Pavlakos
Abstract
In this paper, we present a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often result in physical artifacts such as interpenetration and object floating. To overcome these issues, we introduce a physics-guided reconstruction framework that begins with a kinematic estimate and then refines it through a reinforcement learning (RL) policy trained to reproduce the interaction in a physics simulator. Because kinematic estimates are typically noisy, naive RL training can fail. Therefore, we propose an adaptive sampling strategy with a dual self-updating mechanism that automatically identifies the frames with the most informative and reliable kinematic reconstruction. Our process progressively improves reconstruction quality and yields physically consistent HOI sequences. We demonstrate our approach on two standard benchmarks and achieve clear improvements in physical plausibility metrics over state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0ce1d7a-3ab0-4512-999b-78d9179cb0f0Builds on30
- AMP: adversarial motion priors for stylized physics-based character controlXue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine et al.SIGGRAPH 2021 · 392 citations
- Perpetual Humanoid Control for Real-time Simulated AvatarsZhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani et al.ICCV 2023 · 256 citations
- ASE: large-scale reusable adversarial skill embeddings for physically simulated charactersXue Bin Peng, Yunrong Guo, Lina Halper, Sergey Levine et al.SIGGRAPH 2022 · 217 citations
- BEHAVE: Dataset and Method for Tracking Human Object InteractionsBharat Lal Bhatnagar, Xianghui Xie, Ilya A. Petrov, Cristian Sminchisescu et al.CVPR 2022 · 144 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
Related papers
- Hand-Object Interaction Controller (HOIC): Deep Reinforcement Learning for Reconstructing Interactions with PhysicsHaoyu Hu, Xinyu Yi, Zhe Cao, Jun-Hai Yong et al.SIGGRAPH 2024 · 2 citations
- Contact-guided Real2Sim from Monocular Video with Planar Scene PrimitivesZihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao et al.ICLR 2026
- Stability-driven Contact Reconstruction From Monocular Color ImagesZimeng Zhao, Binghui Zuo, Wei Xie, Yangang WangCVPR 2022 · 15 citations
- Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular VideoErik Gärtner, Mykhaylo Andriluka, Hongyi Xu, Cristian SminchisescuCVPR 2022 · 31 citations
- QuestEnvSim: Environment-Aware Simulated Motion Tracking from Sparse SensorsSunmin Lee, Sebastian Starke, Yuting Ye, Jungdam Won et al.SIGGRAPH 2023 · 31 citations
