What's in your hands? 3D Reconstruction of Generic Objects in Hands
Yufei Ye, Abhinav Gupta, Shubham Tulsiani
Abstract
Our work aims to reconstruct hand-held objects given a single RGB image. In contrast to prior works that typically assume known 3D templates and reduce the problem to 3D pose estimation, our work reconstructs generic hand-held object without knowing their 3D templates. Our key insight is that hand articulation is highly predictive of the object shape, and we propose an approach that conditionally reconstructs the object based on the articulation and the visual input. Given an image depicting a hand-held object, we first use off-the-shelf systems to estimate the underlying hand pose and then infer the object shape in a normalized hand-centric coordinate frame. We parameterized the object by signed distance which are inferred by an implicit network which leverages the information from both visual feature and articulation-aware coordinates to process a query point. We perform experiments across three datasets and show that our method consistently outperforms baselines and is able to reconstruct a diverse set of objects. We analyze the benefits and robustness of explicit articulation conditioning and also show that this allows the hand pose estimation to further improve in test-time optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb018ad8-2192-4812-8294-e67cd1210b8eCited by top-tier papers51
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 80 citations
- Full-Body Articulated Human-Object InteractionNan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui et al.ICCV 2023 · 80 citations
- Omnigrasp: Grasping Diverse Objects with Simulated HumanoidsZhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Winkler et al.NeurIPS 2024 · 66 citations
- Look Ma, No Hands! Agent-Environment Factorization of Egocentric VideosMatthew Chang, Aditya Prakash, Saurabh GuptaNeurIPS 2023 · 26 citations
- Novel-view Synthesis and Pose Estimation for Hand-Object Interaction from Sparse ViewsWentian Qu, Zhaopeng Cui, Yinda Zhang, Chenyu Meng et al.ICCV 2023 · 26 citations
Builds on13
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 184 citations
- CvxNet: Learnable Convex DecompositionBoyang Deng, Kyle Genova, Soroosh Yazdani, Sofien Bouaziz et al.CVPR 2020
- Understanding Human Hands in Contact at Internet ScaleDandan Shan, Jiaqi Geng, Michelle Shu, David F. FouheyCVPR 2020
Related papers
- Learning Explicit Contact for Implicit Reconstruction of Hand-Held Objects from Monocular ImagesJunxing Hu, Hongwen Zhang, Zerui Chen, Mengcheng Li et al.AAAI 2024 · 15 citations
- EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the WildYumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu et al.CVPR 2025
- HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoZicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen et al.CVPR 2024
- Hand-held Object Reconstruction from RGB Video with Dynamic InteractionShijian Jiang, Qi Ye, Rengan Xie, Yuchi Huo et al.CVPR 2025
- gSDF: Geometry-Driven Signed Distance Functions for 3D Hand-Object ReconstructionZerui Chen, Shizhe Chen, Cordelia Schmid, Ivan LaptevCVPR 2023
