Monocular Human-Object Reconstruction in the Wild
Chaofan Huo, Ye Shi, Jingya Wang
摘要
Learning the prior knowledge of the 3D human-object spatial relation is crucial for reconstructing human-object interaction from images and understanding how humans interact with objects in 3D space. Previous works learn this prior from the latest-released human-object interaction dataset collected in controlled environments. However, due to the domain divergence, these methods are limited by the data that the prior learned from and fail to generalize to real-world data with high diversity. To overcome this limitation, we present a 2D-supervised method that learns the 3D humanobject spatial relation prior purely from 2D images in the wild.
Our method utilizes a flow-based neural network to learn the prior distribution of the 2D human-object keypoint layout and viewports for each image in the dataset. The effectiveness of the prior learned from 2D images is demonstrated on the human-object reconstruction task by applying the prior to tune the relative pose between the human and the object during the post-optimization stage. To validate and benchmark our method on in-the-wild images, we collect the WildHOI dataset from the YouTube website, which consists of various interactions with 8 objects in real-world scenarios. We conduct the experiments on the indoor BEHAVE dataset and the outdoor WildHOI dataset. The results show that our method achieves almost comparable performance with fully 3D supervised methods on the BEHAVE dataset, even if we have only utilized the 2D layout information, and outperforms previous methods in terms of generality and interaction diversity on in-the-wild images. The code and the dataset are available at https://huochf.github.io/WildHOI/ for research purposes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni 等NeurIPS 2025 · 被引用 25 次
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti 等CVPR 2026 · 被引用 16 次
- ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object InteractionsZikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li 等CVPR 2026 · 被引用 4 次
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou 等CVPR 2025
- Recovering Physically Plausible Human-Object Interactions from Monocular VideosDingbang Huang, Etienne Vouga, Qixing Huang, Georgios PavlakosCVPR 2026
它引用的顶会 Paper13
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 被引用 1,105 次
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 被引用 482 次
- Deep Mesh Reconstruction From Single RGB Images via Topology Modification NetworksJunyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang 等ICCV 2019 · 被引用 218 次
- Probabilistic Modeling for Human Mesh RecoveryNikos Kolotouros, Georgios Pavlakos, Dinesh Jayaraman, Kostas DaniilidisICCV 2021 · 被引用 201 次
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 被引用 184 次
相关 Paper
- Detailed 2D-3D Joint Representation for Human-Object InteractionYong-Lu Li, Xinpeng Liu, Han Lu, Shiyi Wang 等CVPR 2020
- EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the WildYumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu 等CVPR 2025
- CHORUS: Learning Canonicalized 3D Human-Object Spatial Relations from Unbounded Synthesized ImagesSookwan Han, Hanbyul JooICCV 2023 · 被引用 19 次
- HORP: Human-Object Relation Priors Guided HOI DetectionPei Geng, Jian Yang, Shanshan ZhangCVPR 2025
- CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction ReconstructionPei Geng, Shanshan Zhang, Jian YangCVPR 2026
