TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
Hyeongjin Nam, Daniel Jung, Kyoung Mu Lee
Abstract
Joint reconstruction of 3D human and object from a single image is an active research area, with pivotal applications in robotics and digital content creation. Despite recent advances, existing approaches suffer from two fundamental limitations. First, their reconstructions rely heavily on physical contact information, which inherently cannot capture non-contact human–object interactions, such as gazing at or pointing toward an object. Second, the reconstruction process is primarily driven by local geometric proximity, neglecting the human and object appearances that provide global context crucial for understanding holistic interactions. To address these issues, we introduce TeHOR, a framework built upon two core designs. First, beyond contact information, our framework leverages text descriptions of human–object interactions to enforce semantic alignment between the 3D reconstruction and its textual cues, enabling reasoning over a wider spectrum of interactions, including non-contact cases. Second, we incorporate appearance cues of the 3D human and object into the alignment process to capture holistic contextual information, thereby ensuring visually plausible reconstructions. As a result, our framework produces accurate and semantically coherent reconstructions, achieving state-of-the-art performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- Joint Reconstruction of 3D Human and Object via Contact-Based Refinement TransformerHyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, Kyoung Mu LeeCVPR 2024 · 8 citations
- Human Interaction-Aware 3D Reconstruction from a Single ImageGwanghyun Kim, Junghun James Kim, Suh Yoon Jeon, Jason Park et al.CVPR 2026
- CG-HOI: Contact-Guided 3D Human-Object Interaction GenerationChristian Diller, Angela DaiCVPR 2024
- Stability-driven Contact Reconstruction From Monocular Color ImagesZimeng Zhao, Binghui Zuo, Wei Xie, Yangang WangCVPR 2022 · 15 citations
- PARTE: Part-Guided Texturing for 3D Human Reconstruction from a Single ImageHyeongjin Nam, Donghwan Kim, Gyeongsik Moon, Kyoung Mu LeeICCV 2025 · 1 citation
