MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation
Bohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing Lu
Abstract
Egocentric hand-object motion generation is crucial for immersive AR/VR and robotic imitation but remains challenging due to unstable viewpoints, selfocclusions, perspective distortion, and noisy ego-motion. Existing methods rely on predefined 3D object priors, limiting generalization to novel objects, which restricts their generalizability to novel objects. Meanwhile, recent multimodal approaches suffer from ambiguous generation from abstract textual cues, intricate pipelines for modeling 3D hand-object correlation, and compounding errors in open-loop prediction. We propose MEgoHand, a multimodal framework that synthesizes physically plausible hand-object interactions from egocentric RGB, text, and initial hand pose. MEgoHand introduces a bi-level architecture: a high-level "cerebrum" leverages a vision language model (VLM) to infer motion priors from visual-textual context and a monocular depth estimator for object-agnostic spatial reasoning, while a low-level DiT-based flow-matching policy generates fine-grained trajectories with temporal orthogonal filtering to enhance stability. To address dataset inconsistency, we design a dataset curation paradigm with an Inverse MANO Retargeting Network and Virtual RGB-D Renderer, curating a unified dataset of 3.35M RGB-D frames, 24K interactions, and 1.2K objects. Extensive experiments across five in-domain and two cross-domain datasets demonstrate the effectiveness of MEgo-Hand, achieving substantial reductions in wrist translation error (86.9%) and joint rotation error (34.1%), highlighting its capacity to accurately model fine-grained hand joint structures and generalize robustly across diverse scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b9b0fcb-e439-4518-960e-357cd7e961c0Cited by top-tier papers2
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware RepresentationHaodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan et al.CVPR 2026 · 5 citations
- PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video GenerationMingju Gao, Kaisen Yang, Huan-ang Gao, Bohan Li et al.CVPR 2026 · 3 citations
Builds on23
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
- HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real WorldXin Wang, Taein Kwon, Mahdi Rad, Bowen Pan et al.ICCV 2023 · 151 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa et al.CVPR 2024 · 110 citations
Related papers
- EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context LearningBinzhu Xie, Shi Qiu, Sicheng Zhang, Yinqiao Wang et al.ICLR 2026 · 4 citations
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu et al.ICLR 2026 · 11 citations
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni et al.NeurIPS 2025 · 25 citations
- AGILE: Hand-object Interaction Reconstruction from Video via Agentic GenerationJin-Chuan Shi, Binhong Ye, Tao Liu, Xiaoyang Liu et al.SIGGRAPH 2026
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri et al.CVPR 2025
