Contact-guided Real2Sim from Monocular Video with Planar Scene Primitives
Zihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao, Jessica K. Hodgins, Shubham Tulsiani, Deva Ramanan
Abstract
We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human--scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or recovers noisy geometry with artifacts that cause motion-tracking policies with scene interactions to fail. In contrast, our key insight is to fit simulation-ready convex planar primitives to a depth-based point cloud reconstruction of the scene via a simple clustering pipeline over depth, normals, and flow. To reconstruct scene geometry that might be occluded during interactions, we use human--scene contact modeling (e.g., using human posture to reconstruct the occluded seat of a chair). Finally, we ensure that human and scene reconstructions are physically plausible by using them to drive a humanoid controller via reinforcement learning. Our approach reduces motion-tracking failure rates from 55.2% to 6.9% on human-centric video benchmarks (EMDB, PROX), while delivering 43% faster RL simulation throughput. This demonstrates CRISP's ability to generate physically valid human motion and interaction environments at scale, advancing real-to-sim applications for robotics. Code and interactive demos are available at our project website: https://crisp-real2sim.github.io/CRISP-Real2Sim
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- OnlineHMR: Video-based Online World-Grounded Human Mesh RecoveryYiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang et al.CVPR 2026 · 5 citations
- EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied AgentsWenjia Wang, Liang Pan, Huaijin Pi, Yuke Lou et al.CVPR 2026 · 2 citations
Builds on25
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat et al.ICCV 2023 · 414 citations
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 384 citations
- Perpetual Humanoid Control for Real-time Simulated AvatarsZhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani et al.ICCV 2023 · 256 citations
- Unified Human-Scene Interaction via Prompted Chain-of-ContactsZeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao et al.ICLR 2024 · 113 citations
Related papers
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- Hand-Object Interaction Controller (HOIC): Deep Reinforcement Learning for Reconstructing Interactions with PhysicsHaoyu Hu, Xinyu Yi, Zhe Cao, Jun-Hai Yong et al.SIGGRAPH 2024 · 2 citations
- Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular VideoErik Gärtner, Mykhaylo Andriluka, Hongyi Xu, Cristian SminchisescuCVPR 2022 · 31 citations
- QuestEnvSim: Environment-Aware Simulated Motion Tracking from Sparse SensorsSunmin Lee, Sebastian Starke, Yuting Ye, Jungdam Won et al.SIGGRAPH 2023 · 31 citations
- Human-Aware Object Placement for Visual Environment ReconstructionHongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas et al.CVPR 2022 · 61 citations
