HOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data
Mengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu, Zhuowen Tu, Xiaolong Wang
Abstract
3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection pro-cess. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction data. Our model is a conditional diffusion model that takes both the 3D hand-object geometric structure and text description as inputs for image synthesis. This offers a more control-lable and realistic synthesis as we can specify the structure and style inputs in a disentangled manner. HOIDiffusion is trained by leveraging a diffusion model pre-trained on large-scale natural images and a few 3D human demonstrations. Beyond controllable image synthesis, we adopt the generated 3D data for learning 6D object pose estimation and show its effectiveness in improving perception systems. Project page: https://mq-zhang1.github.io/HOIDiffusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79e8a01-ec25-497e-a73c-c47a9fcc7854Cited by top-tier papers16
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee et al.CVPR 2026 · 13 citations
- EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric ObservationsJunho Park, Andrew Sangwoo Ye, Taein KwonICLR 2026 · 10 citations
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang et al.CVPR 2026 · 5 citations
- What Are You Doing? A Closer Look at Controllable Human Video GenerationEmanuele Bugliarello, Anurag Arnab, Roni Paiss, Christy Koh et al.CVPR 2026 · 4 citations
- HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data AugmentationWentian Qu, Jiahe Li, Jian Cheng, Jian Shi et al.AAAI 2025 · 4 citations
Builds on37
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
Related papers
- HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object InteractionsHao Xu, Haipeng Li, Yinqiao Wang, Shuaicheng Liu et al.CVPR 2024 · 12 citations
- HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point CloudWencan Cheng, Hao Tang, Luc Van Gool, Jong Hwan KoCVPR 2024
- Template Free Reconstruction of Human-object Interaction with Procedural Interaction GenerationXianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, Gerard Pons-MollCVPR 2024 · 6 citations
- Single-view Image to Novel-view Generation for Hand-Object InteractionsZhongqun Zhang, Yihua Cheng, Eduardo Pérez-Pellitero, Yiren Zhou et al.AAAI 2025 · 1 citation
- HandDiffuse: Generative Controllers for Two-Hand Interactions via Diffusion ModelsPei LinAAAI 2025 · 1 citation
