Joint Optimization for 4D Human-Scene Reconstruction in the Wild
Zhizheng Liu, Joe Lin, Wayne Wu, Bolei Zhou
摘要
Reconstructing human motion and its surrounding environment is crucial for understanding human-scene interaction and predicting human movements in the scene. While much progress has been made in capturing human-scene interaction in constrained environments, those prior methods can hardly reconstruct the natural and diverse human motion and scene context from web videos. In this work, we propose JOSH, a novel optimization-based method for 4D human-scene reconstruction in the wild from monocular videos. Compared to prior works that perform separate optimization of the human, the camera, and the scene, JOSH leverages the human-scene contact constraints to jointly optimize all parameters in a single stage. Experiment results demonstrate that JOSH significantly improves 4D human-scene reconstruction, global human motion estimation, and dense scene reconstruction by utilizing the joint optimization of scene geometry, human motion, and camera poses. Further studies show that JOSH can enable scalable training of end-to-end global human motion models on extensive web data, highlighting its robustness and generalizability. The code and model are available at https://vail-ucla.github.io/JOSH/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Human3R: Everyone Everywhere All at OnceYue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen 等ICLR 2026 · 被引用 38 次
- UniSH: Unifying Scene and Human Reconstruction in a Feed-Forward PassMengfei Li, Peng Li, Zheng Zhang, Jiahao Lu 等CVPR 2026 · 被引用 7 次
- DuoMo: Dual Motion Diffusion for World-Space Human ReconstructionYufu Wang, Evonne Ng, Soyong Shin, Rawal Khirodkar 等CVPR 2026 · 被引用 6 次
- OnlineHMR: Video-based Online World-Grounded Human Mesh RecoveryYiwen Zhao, Ce Zheng, Yufu Wang, Hsueh-Han Daniel Yang 等CVPR 2026 · 被引用 5 次
- HAMSt3R: Human-Aware Multi-View Stereo 3D ReconstructionSara Rojas, Matthieu Armando, Bernard Ghanem, Philippe Weinzaepfel 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper25
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- 4D Gaussian Splatting for Real-Time Dynamic Scene RenderingGuanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie 等CVPR 2024 · 被引用 513 次
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang 等ICCV 2021 · 被引用 398 次
相关 Paper
- C4D: 4D Made from 3D Through Dual CorrespondencesShizun Wang, Zhenxiang Jiang, Xingyi Yang, Xinchao WangICCV 2025 · 被引用 4 次
- Uni4D: Unifying Visual Foundation Models for 4D Modeling from a Single VideoDavid Yifan Yao, Albert J. Zhai, Shenlong WangCVPR 2025
- Human-Aware Object Placement for Visual Environment ReconstructionHongwei Yi, Chun-Hao P. Huang, Dimitrios Tzionas, Muhammed Kocabas 等CVPR 2022 · 被引用 61 次
- Crowd4D: Scene-Aware Monocular 4D Crowd ReconstructionHongbo Kang, Tianyi Zhou, Qingyang Yang, Hongwei wen 等ICML 2026
- Learning Motion Priors for 4D Human Body Capture in 3D ScenesSiwei Zhang, Yan Zhang, Federica Bogo, Marc Pollefeys 等ICCV 2021 · 被引用 117 次
