Leveraging Photometric Consistency Over Time for Sparsely Supervised Hand-Object Reconstruction
Yana Hasson, Bugra Tekin, Federica Bogo, Ivan Laptev, Marc Pollefeys, Cordelia Schmid
Abstract
Modeling hand-object manipulations is essential for understanding how humans interact with their environment. While of practical importance, estimating the pose of hands and objects during interactions is challenging due to the large mutual occlusions that occur during manipulation. Recent efforts have been directed towards fully-supervised methods that require large amounts of labeled training samples. Collecting 3D ground-truth data for hand-object interactions, however, is costly, tedious, and error-prone. To overcome this challenge we present a method to leverage photometric consistency across time when annotations are only available for a sparse subset of frames in a video. Our model is trained end-to-end on color images to jointly reconstruct hands and objects in 3D by inferring their poses. Given our estimated reconstructions, we differentiably render the optical flow between pairs of adjacent images and use it within the network to warp one frame to another. We then apply a self-supervised photometric loss that relies on the visual consistency between nearby images. We achieve state-of-the-art results on 3D hand-object reconstruction benchmarks and demonstrate that our approach allows us to improve the pose estimation accuracy by leveraging information from neighboring frames in low-data regimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers77
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 184 citations
- CPF: Learning a Contact Potential Field to Model the Hand-Object InteractionLixin Yang, Xinyu Zhan, Kailin Li, Wenqiang Xu et al.ICCV 2021 · 170 citations
- HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation NetworkJoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi et al.CVPR 2022 · 116 citations
- Interacting Two-Hand 3D Pose and Shape Reconstruction from Single Color ImageBaowen Zhang, Yangang Wang, Xiaoming Deng, Yinda Zhang et al.ICCV 2021 · 114 citations
Builds on5
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell et al.ICCV 2019 · 493 citations
- End-to-End Hand Mesh Recovery From a Monocular RGB ImageXiong Zhang, Qiang Li, Hong Mo, Wenbo Zhang et al.ICCV 2019 · 248 citations
- TexturePose: Supervising Human Mesh Estimation With Texture ConsistencyGeorgios Pavlakos, Nikos Kolotouros, Kostas DaniilidisICCV 2019 · 109 citations
- Mono-SF: Multi-View Geometry Meets Single-View Depth for Monocular Scene Flow Estimation of Dynamic Traffic ScenesFabian Brickwedde, Steffen Abraham, Rudolf MesterICCV 2019 · 55 citations
Related papers
- Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in TimeShaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu et al.CVPR 2021
- Model-Based 3D Hand Reconstruction via Self-Supervised LearningYujin Chen, Zhigang Tu, Di Kang, Linchao Bao et al.CVPR 2021
- HOnnotate: A Method for 3D Annotation of Hand and Object PosesShreyas Hampali, Mahdi Rad, Markus Oberweger, Vincent LepetitCVPR 2020
- MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video ClipsShibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt et al.ICCV 2025 · 2 citations
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu et al.CVPR 2020
