Single-view Image to Novel-view Generation for Hand-Object Interactions
Zhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou, Jiankang Deng, Hyung Jin Chang, Jifei Song
Abstract
Hand-object interaction modeling from a single RGB image is a significantly challenging task. Previous works typically reconstruct hand-object interactions as texture-less meshes, ignoring photo-realistic image generation. In this work, we introduce the HO123, a novel method to synthesize novel-view hand-object interaction images from a single image. To this end, we first train a 2D diffusion prior. Given the camera pose in novel views, our approach transfers the camera information into explicit hand representations, including hand depth and skeleton images. We propose a global hand embedding to control the diffusion model based on these hand representations. We then learn a 3D Gaussian splatting for novel-view rendering using the diffusion prior. However, occluded objects present a persistent challenge. To address this issue, we further introduce local hand embedding, where a contact field is defined in the 3D Gaussian Splatting. We leverage contact information to guide the rendering in the contact field. Extensive experiments on the HO3D and DexYCB datasets demonstrate that our method significantly outperforms state-of-the-art novel-view synthesis for hand-object interactions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57e682e7-a3c2-4cfa-85a7-9db9bcb9f4e2Cited by top-tier papers2
- Force-Aware 3D Contact Modeling for Stable Grasp GenerationZhuo Chen, Zhongqun Zhang, Yihua Cheng, Ales Leonardis et al.AAAI 2026
- MGD: Mesh-guided Gaussians with Diffusion Priors for Dynamic Objects Reconstruction from Monocular RGB-D VideoWeixing Xie, Ying Ye, Xian Wu, Jintian Li et al.AAAI 2026
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov et al.ICCV 2023 · 1,662 citations
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
Related papers
- Novel-view Synthesis and Pose Estimation for Hand-Object Interaction from Sparse ViewsWentian Qu, Zhaopeng Cui, Yinda Zhang, Chenyu Meng et al.ICCV 2023 · 26 citations
- Generalizable Hand-Object Modeling from Monocular RGB Images via 3D GaussiansXingyu Liu, Pengfei Ren, Qi Qi, Haifeng Sun et al.NeurIPS 2025 · 5 citations
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu et al.CVPR 2024 · 12 citations
- Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand AvatarsXuan Huang, Hanhui Li, Wanquan Liu, Xiaodan Liang et al.NeurIPS 2024 · 6 citations
- Hand-Object Interaction Image GenerationHezhen Hu, Weilun Wang, Wengang Zhou, Houqiang LiNeurIPS 2022 · 24 citations
