CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction
Pei Geng, Shanshan Zhang, Jian Yang
Abstract
Reconstructing 3D human-object interaction (HOI) from monocular images is highly challenging especially when human and object are mutually occluded. Existing methods primarily rely on single-view inputs, which fundamentally limit their ability to recover occluded regions and accurately estimate contact areas. To address these challenges, we for the first time, consider to introduce novelview feature priors to enhance monocular 3D HOI reconstruction. We first design a cross-view generator that learns to infer novel-view image features from a single-view input, enriching spatial geometry at the feature level without requiring extra inputs during inference. Guided by both real and generated view features, a spatial crossview feature fusion module adaptively aggregates complementary cues to enhance the initial reconstruction of human and object meshes. Built upon this reconstruction, we sample 3D vertex features from both views and introduce a bidirectional cross-view Transformer to integrate multi-view vertex representations for accurate contact estimation. Finally, the predicted contact maps are leveraged to refine human-object meshes, yielding geometrically consistent and physically plausible reconstructions. Experiments on BEHAVE and InterCap show that our proposed CrossHOI surpasses state-of-the-art methods in both reconstruction accuracy and contact prediction, especially under severe occlusions. Code is available at https: //github.com/peigeng99/CrossHOI.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6f7157ca-2cdd-4501-afed-1f0b814510faBuilds on22
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 399 citations
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu et al.NeurIPS 2023 · 267 citations
Related papers
- End-to-End HOI Reconstruction Transformer with Graph-based EncodingZhenrong Wang, Qi Zheng, Sihan Ma, Maosheng Ye et al.CVPR 2025
- Reconstructing In-the-Wild Open-Vocabulary Human-Object InteractionsBoran Wen, Dingbang Huang, Zichen Zhang, Jiahong Zhou et al.CVPR 2025
- Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesJunxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li et al.AAAI 2024 · 15 citations
- HORP: Human-Object Relation Priors Guided HOI DetectionPei Geng, Jian Yang, Shanshan ZhangCVPR 2025
- Detailed 2D-3D Joint Representation for Human-Object InteractionYong-Lu Li, Xinpeng Liu, Han Lu, Shiyi Wang et al.CVPR 2020
