Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan, Xin Gong, Zehang Luo, Chengxi Heyu, Junfeng Li, Wenxuan Song, Shunbo Zhou, Haoang Li
摘要
Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current methods utilize HOI representation as an auxiliary generative objective to guide video synthesis. However, there is a dilemma between 2D and 3D representations that cannot simultaneously guarantee scalability and interaction fidelity. To address this limitation, we propose a structure and contact-aware representation that captures hand-object contact, hand-object occlusion, and holistic structure context without 3D annotations. This interaction-oriented and scalable supervision signal enables the model to learn fine-grained interaction physics and generalize to open-world scenarios. To fully exploit the proposed representation, we introduce a joint-generation paradigm with a share-and-specialization strategy that generates interaction-oriented representations and videos. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on two real-world datasets in generating physics-realistic and temporally coherent HOI videos. Furthermore, our approach exhibits strong generalization to challenging open-world scenarios, highlighting the benefit of our scalable design. Our project page is https://hgzn258.github.io/SCAR/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction ScenariosLingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min 等NeurIPS 2025 · 被引用 12 次
- Hand-Object Interaction Image GenerationHezhen Hu, Weilun Wang, Wengang Zhou, Houqiang LiNeurIPS 2022 · 被引用 24 次
- ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction VideosYuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang 等CVPR 2026 · 被引用 6 次
- HOLD: Category-Agnostic 3D Reconstruction of Interacting Hands and Objects from VideoZicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen 等CVPR 2024
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee 等CVPR 2026 · 被引用 13 次
