MagicHOI: Leveraging 3D Priors for Accurate Hand-Object Reconstruction from Short Monocular Video Clips
Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt, Zicong Fan, Jie Song
摘要
Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera viewpoints and static grips leave parts of the object unobserved, resulting in implausible reconstructions. To overcome this, we present MagicHOI, a method for reconstructing hands and objects from short monocular interaction videos, even under limited viewpoint variation. Our key insight is that, despite the scarcity of paired 3D hand-object data, largescale novel view synthesis diffusion models offer rich object supervision. This supervision serves as a prior to regularize unseen object regions during hand interactions. Leveraging this insight, we integrate a novel view synthesis model into our hand-object reconstruction framework. We further align hand to object by incorporating visible contact constraints. Our results demonstrate that MagicHOI significantly outperforms existing state-of-the-art hand-object reconstruction methods. We also show that novel view synthesis diffusion priors effectively regularize unseen object regions, enhancing 3D hand-object reconstruction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction VideosYuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang 等CVPR 2026 · 被引用 6 次
- ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object InteractionsZikai Wang, Zhilu Zhang, Yiqing Wang, Hui Li 等CVPR 2026 · 被引用 4 次
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li 等CVPR 2026 · 被引用 3 次
它引用的顶会 Paper44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Volume Rendering of Neural Implicit SurfacesLior Yariv, Jiatao Gu, Yoni Kasten, Yaron LipmanNeurIPS 2021 · 被引用 1,421 次
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai 等ICLR 2024 · 被引用 973 次
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi 等ICLR 2024 · 被引用 813 次
相关 Paper
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 被引用 80 次
- HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoZicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen 等CVPR 2024
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction ScenariosLingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min 等NeurIPS 2025 · 被引用 12 次
- Diffusion-Based 3D Hand Motion Recovery with Intuitive PhysicsYufei Zhang, Zijun Cui, Jeffrey O. Kephart, Qiang JiICCV 2025 · 被引用 1 次
- EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the WildYumeng Liu, Xiaoxiao Long, Zemin Yang, Yuan Liu 等CVPR 2025
