Contact-guided Real2Sim from Monocular Video with Planar Scene Primitives
Zihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao, Jessica K. Hodgins, Shubham Tulsiani, Deva Ramanan
摘要
We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human--scene reconstruction relies on data-driven priors and joint optimization with no physics in the loop, or recovers noisy geometry with artifacts that cause motion-tracking policies with scene interactions to fail. In contrast, our key insight is to fit simulation-ready convex planar primitives to a depth-based point cloud reconstruction of the scene via a simple clustering pipeline over depth, normals, and flow. To reconstruct scene geometry that might be occluded during interactions, we use human--scene contact modeling (e.g., using human posture to reconstruct the occluded seat of a chair). Finally, we ensure that human and scene reconstructions are physically plausible by using them to drive a humanoid controller via reinforcement learning. Our approach reduces motion-tracking failure rates from 55.2% to 6.9% on human-centric video benchmarks (EMDB, PROX), while delivering 43% faster RL simulation throughput. This demonstrates CRISP's ability to generate physically valid human motion and interaction environments at scale, advancing real-to-sim applications for robotics. Code and interactive demos are available at our project website: https://crisp-real2sim.github.io/CRISP-Real2Sim
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OnlineHMR: Video-based Online World-Grounded Human Mesh RecoveryYiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang 等CVPR 2026 · 被引用 5 次
- EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied AgentsWenjia Wang, Liang Pan, Huaijin Pi, Yuke Lou 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper25
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat 等ICCV 2023 · 被引用 414 次
- Resolving 3D Human Pose Ambiguities With 3D Scene ConstraintsMohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, Michael J. BlackICCV 2019 · 被引用 384 次
- Perpetual Humanoid Control for Real-time Simulated AvatarsZhengyi Luo, Jinkun Cao, Alexander Winkler, Kris Kitani 等ICCV 2023 · 被引用 256 次
- Unified Human-Scene Interaction via Prompted Chain-of-ContactsZeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao 等ICLR 2024 · 被引用 113 次
相关 Paper
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
- Hand-Object Interaction Controller (HOIC): Deep Reinforcement Learning for Reconstructing Interactions with PhysicsHaoyu Hu, Xinyu Yi, Zhe Cao, Jun-Hai Yong 等SIGGRAPH 2024 · 被引用 2 次
- Trajectory Optimization for Physics-Based Reconstruction of 3d Human Pose from Monocular VideoErik Gärtner, Mykhaylo Andriluka, Hongyi Xu, Cristian SminchisescuCVPR 2022 · 被引用 31 次
- QuestEnvSim: Environment-Aware Simulated Motion Tracking from Sparse SensorsSunmin Lee, Sebastian Starke, Yuting Ye, Jungdam Won 等SIGGRAPH 2023 · 被引用 31 次
- Human-Aware Object Placement for Visual Environment ReconstructionHongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas 等CVPR 2022 · 被引用 61 次
