ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
Youxin Pang, Ruizhi Shao, Jiajun Zhang, Hanzhang Tu, Yun Liu, Boyao Zhou, Hongwen Zhang, Yebin Liu
Abstract
Occlusion-free normal maps Occlusion confidence maps Confidence occlusion occlusion-free Unseen objects Generated results Human-centered results Segmentation Figure 1. ManiVideo: We propose a novel framework for generalizable and dexterous hand-object manipulation video generation. Left: Given several reference images of unseen objects, our method generates realistic and plausible manipulation videos driven by hand-object signals. By integrating multiple datasets, ManiVideo supports applications such as human-centered manipulation video generation. Right:
To ensure hand-object consistency, we introduce a multi-layer occlusion representation capable of learning 3D occlusion relationships from occlusion-free normal maps and occlusion confidence maps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2cfb0e9b-0e8a-45d5-abe1-c65940f0cff4Cited by top-tier papers1
Ask how each one uses itBuilds on37
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 242 citations
- Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion ModelsFei Shen, Hu Ye, Jun Zhang, Cong Wang et al.ICLR 2024 · 133 citations
Related papers
- Diffusion-Guided Reconstruction of Everyday Hand-Object Interaction ClipsYufei Ye, Poorvi Hebbar, Abhinav Gupta, Shubham TulsianiICCV 2023 · 80 citations
- InterDyn: Controllable Interactive Dynamics with Video Diffusion ModelsRick Akkerman, Haiwen Feng, Michael J. Black, Dimitrios Tzionas et al.CVPR 2025
- Hand Avatar: Free-Pose Hand Animation and Rendering from Monocular VideoXingyu Chen, Baoyuan Wang, Heung-Yeung ShumCVPR 2023
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang et al.ICLR 2026
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 6 citations
