Joint Reconstruction of 3D Human and Object via Contact-Based Refinement Transformer
Hyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, Kyoung Mu Lee
Abstract
Human-object contact serves as a strong cue to understand how humans physically interact with objects. Nev-ertheless, it is not widely explored to utilize human-object contact information for the joint reconstruction of 3D human and object from a single image. In this work, we present a novel joint 3D human-object reconstruction method (CONTHO) that effectively exploits contact information between humans and objects. There are two core designs in our system: 1) 3D-guided contact estimation and 2) contact-based 3D human and object refinement. First, for accurate human-object contact estimation, CONTHO initially reconstructs 3D humans and objects and utilizes them as explicit 3D guidance for contact estimation. Second, to refine the initial reconstructions of 3D human and object, we propose a novel contact-based refinement Transformer that effectively aggregates human features and object features based on the estimated human-object contact. The proposed contact-based refinement prevents the learning of erroneous correlation between human and object, which enables accurate 3D reconstruction. As a result, our CON-THO achieves state-of-the-art performance in both human-object contact estimation and joint reconstruction of 3D human and object. The code is publicly available<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup><sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>https://github.com/dqj5182/CONTHO_RELEASE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bf6b2f6-746a-42a8-9a21-3a59c52d074eCited by top-tier papers19
- EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric ViewsYuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu et al.NeurIPS 2024 · 31 citations
- CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial IntelligenceJingyu Shi, Rahul Jain, Seunggeun Chi, Hyungjun Doh et al.CHI 2025 · 19 citations
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti et al.CVPR 2026 · 16 citations
- Learning Dense Hand Contact Estimation from Imbalanced DataDaniel Sungho Jung, Kyoung Mu LeeNeurIPS 2025 · 14 citations
- Humoto: A 4D Dataset of Mocap Human Object InteractionsJiaxin Lu, Chun-Hao Paul Huang, Uttaran Bhattacharya, Qixing Huang et al.ICCV 2025 · 4 citations
Builds on22
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- HuMoR: 3D Human Motion Model for Robust Pose EstimationDavis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang et al.ICCV 2021 · 398 citations
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback LoopHongwen Zhang, Yating Tian, Xinchi Zhou, Wanli Ouyang et al.ICCV 2021 · 376 citations
- ZebraPose: Coarse to Fine Surface Encoding for 6DoF Object Pose EstimationYongzhi Su, Mahdi Saleh, Torben Fetzer, Jason R. Rambach et al.CVPR 2022 · 170 citations
Related papers
- TeHOR: Text-Guided 3D Human and Object Reconstruction with TexturesHyeongjin Nam, Daniel Jung, Kyoung Mu LeeCVPR 2026 · 1 citation
- ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided DiffusionAo Li, Jinpeng Liu, Yixuan Zhu, Yansong TangICCV 2025 · 1 citation
- Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesJunxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li et al.AAAI 2024 · 15 citations
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
- CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction ReconstructionPei Geng, Shanshan Zhang, Jian YangCVPR 2026
