G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis
Yufei Ye, Abhinav Gupta, Kris Kitani, Shubham Tulsiani
摘要
We propose G-HOP, a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand, conditioned on the object category. To learn a 3D spatial diffusion model that can capture this joint distribution, we represent the human hand via a skeletal distance field to obtain a representation aligned with the (latent) signed distance field for the object. We show that this hand-object prior can then serve as generic guidance to facilitate other tasks like reconstruction from interaction clip and human grasp synthesis. We believe that our model, trained by aggregating seven diverse real-world interaction datasets spanning across 155 cate-gories, represents a first approach that allows jointly generating both hand and object. Our empirical evaluations demonstrate the benefit of this joint prior in video-based reconstruction and human grasp synthesis, outperforming current task-specific baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Omnigrasp: Grasping Diverse Objects with Simulated HumanoidsZhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Winkler 等NeurIPS 2024 · 被引用 66 次
- CARI4D: Category Agnostic 4D Reconstruction of Human-Object InteractionXianghui Xie, Bowen Wen, Yan Chang, Hesam Rabeti 等CVPR 2026 · 被引用 16 次
- MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion GenerationBohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing LuNeurIPS 2025 · 被引用 14 次
- TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object InteractionsGuangyi Han, Wei Zhai, Yuhang Yang, Yang Cao 等ICLR 2026 · 被引用 11 次
- ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction VideosYuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-1-to-3: Zero-shot One Image to 3D ObjectRuoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov 等ICCV 2023 · 被引用 1,662 次
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 被引用 463 次
相关 Paper
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 被引用 80 次
- Single-view Image to Novel-view Generation for Hand-Object InteractionsZhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou 等AAAI 2025 · 被引用 1 次
- LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand DiffusionMuchen Li, Sammy Christen, Chengde Wan, Yujun Cai 等CVPR 2025
- HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance FieldsHaozhe Qi, Chen Zhao, Mathieu Salzmann, Alexander MathisCVPR 2024 · 被引用 18 次
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee 等CVPR 2026 · 被引用 13 次
