LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand Diffusion
Muchen Li, Sammy Christen, Chengde Wan, Yujun Cai, Renjie Liao, Leonid Sigal, Shugao Ma
Abstract
Current research on generating 3D hand-object interaction motion primarily focuses on in-domain objects. Generalization to unseen objects is essential for practical applications, yet it remains both challenging and largely unexplored. In this paper, we propose LatentHOI, a novel approach designed to tackle the challenges of generalizing hand-object interaction synthesis to unseen objects. Our main insight lies in decoupling high-level temporal motion from fine-grained spatial hand-object interactions via a latent diffusion model coupled with a Grasping Variational Autoencoder (Grasp-VAE). This configuration introduces regularization by enforcing a conditional dependency between spatial grasping and temporal motion, as well as through the regularized latent space for better generalization ability. We conducted extensive experiments in an unseen-object setting on both single-hand grasping and bi-manual motion datasets, including GRAB, DexYCB†, and OakInk. Quantitative and qualitative evaluations demonstrate that our method significantly enhances the realism and physical plausibility of generated motions for unseen objects, both in single and bimanual manipulations, compared to the state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14aa6b57-df1d-4482-b574-dac676cf0668Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
Related papers
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction ScenariosLingwei Dang, Ruizhi Shao, Hongwen Zhang, Wei Min et al.NeurIPS 2025 · 12 citations
- Disentangled Hierarchical VAE for 3D Human-Human Interaction GenerationZichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu et al.ICLR 2026 · 3 citations
- MoEG-HOI: Mixture of Expert Groups for One-Stage Hand-Object Interaction Motion Generation with Hand-Finger-Joint Semantic GuidanceHang Xu, Yang Xiao, Changlong Jiang, Haohong Kuang et al.AAAI 2026
- G-DexGrasp: Generalizable Dexterous Grasping Synthesis via Part-Aware Prior Retrieval and Prior-Assisted GenerationJuntao Jian, Xiuping Liu, Zixuan Chen, Manyi Li et al.ICCV 2025 · 2 citations
- HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video SynthesisMingjin Chen, Junhao Chen, Zhaoxin Fan, Yujian Lee et al.CVPR 2026 · 13 citations
