Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular Images
Junxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li, Yunlong Wang, Yebin Liu, Zhenan Sun
摘要
Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent works have employed implicit functions to achieve impressive progress, they ignore formulating contacts in their frameworks, which results in producing less realistic object meshes. In this work, we explore how to model contacts in an explicit way to benefit the implicit reconstruction of hand-held objects. Our method consists of two components: explicit contact prediction and implicit shape reconstruction. In the first part, we propose a new subtask of directly estimating 3D hand-object contacts from a single image. The part-level and vertex-level graph-based transformers are cascaded and jointly learned in a coarse-to-fine manner for more accurate contact probabilities. In the second part, we introduce a novel method to diffuse estimated contact states from the hand mesh surface to nearby 3D space and leverage diffused contact probabilities to construct the implicit neural representation for the manipulated object. Benefiting from estimating the interaction patterns between the hand and the object, our method can reconstruct more realistic object meshes, especially for object parts that are in contact with hands. Extensive experiments on challenging benchmarks show that the proposed method outperforms the current state of the arts by a great margin. Our code is publicly available at https://junxinghu.github.io/projects/hoi.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Dense Hand Contact Estimation from Imbalanced DataDaniel Sungho Jung, Kyoung Mu LeeNeurIPS 2025 · 被引用 14 次
- HORT: Monocular Hand-held Objects Reconstruction with TransformersZerui Chen, Rolandos Alexandros Potamias, Shizhe Chen, Cordelia SchmidICCV 2025 · 被引用 4 次
- End-to-End HOI Reconstruction Transformer with Graph-based EncodingZhenrong Wang, Qi Zheng, Sihan Ma, Maosheng Ye 等CVPR 2025
- VPHO: Joint Visual-Physical Cue Learning and Aggregation for Hand-Object Pose EstimationJun Zhou, Chi Xu, Kaifeng Tang, Yuting Ge 等AAAI 2026
它引用的顶会 Paper18
- Mesh GraphormerKevin Lin, Lijuan Wang, Zicheng LiuICCV 2021 · 被引用 399 次
- Neural Human Performer: Learning Generalizable Radiance Fields for Human Performance RenderingYoungjoong Kwon, Dahun Kim, Duygu Ceylan, Henry FuchsNeurIPS 2021 · 被引用 224 次
- CPF: Learning a Contact Potential Field to Model the Hand-Object InteractionLixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu 等ICCV 2021 · 被引用 170 次
- LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh RegistrationBharat Lal Bhatnagar, Cristian Sminchisescu, Christian Theobalt, Gerard Pons-MollNeurIPS 2020 · 被引用 159 次
- Keypoint Transformer: Solving Joint Identification in Challenging Hands and Object Interactions for Accurate 3D Pose EstimationShreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, Vincent LepetitCVPR 2022 · 被引用 155 次
相关 Paper
- In-Hand 3D Object Reconstruction from a Monocular RGB VideoShijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo 等AAAI 2024 · 被引用 10 次
- What's in your hands? 3D Reconstruction of Generic Objects in HandsYufei Ye, Abhinav Gupta, Shubham TulsianiCVPR 2022 · 被引用 69 次
- HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance FieldsHaozhe Qi, Chen Zhao, Mathieu Salzmann, Alexander MathisCVPR 2024 · 被引用 18 次
- HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoZicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen 等CVPR 2024
- CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction ReconstructionPei Geng, Shanshan Zhang, Jian YangCVPR 2026
