GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and Imitation
Zifan Wang, Junyu Chen, Ziqing Chen, Pengwei Xie, Rui Chen, Li Yi
Abstract
This paper presents GenH2R, a framework for learning generalizable vision-based human-to-robot (H2R) handover skills. The goal is to equip robots with the ability to reliably receive objects with unseen geometry handed over by humans in various complex trajectories. We acquire such generalizability by learning H2R handover at scale with a comprehensive solution including procedural simulation assets creation, automated demonstration generation, and effective imitation learning. We leverage large-scale 3D model repositories, dexterous grasp generation methods, and curve-based 3D animation to create an H2R handover simulation environment named GenH2R-Sim, surpassing the number of scenes in existing simulators by three orders of magnitude. We further introduce a distillation-friendly demonstration generation method that automati-cally generates a million high-quality demonstrations suitable for learning. Finally, we present a 4D imitation learning method augmented by a future forecasting objective to distill demonstrations into a visuo-motor handover policy. Experimental evaluations in both simulators and the real world demonstrate significant improvements (at least +10% success rate) over baselines in all cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3faf269d-bec0-40d3-9d7c-c9f597e847ccCited by top-tier papers6
- UniHM: Unified Dexterous Hand Manipulation with Vision Language ModelZhenhao Zhang, Jiaxin Liu, Ye Shi, Jingya WangICLR 2026 · 4 citations
- Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human InteractionsLiang Xu, Chengqun Yang, Zili Lin, Fei Xu et al.ICCV 2025 · 2 citations
- DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-To-Robot HandoverYouzhuo Wang, Jiayi Ye, Chuyang Xiao, Yiming Zhong et al.ICCV 2025 · 1 citation
- SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction SynthesisWenkun He, Yun Liu, Ruitao Liu, Li YiICCV 2025 · 1 citation
- MobileH2R: Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic DataZifan Wang, Ziqing Chen, Junyu Chen, Jilong Wang et al.CVPR 2025
Builds on9
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- RoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationYufei Wang, Zhou Xian, Feng Chen, Tsun-Hsuan Wang et al.ICML 2024 · 227 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
- DexYCB: A Benchmark for Capturing Hand Grasping of ObjectsYu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov et al.CVPR 2021
Related papers
- Omnigrasp: Grasping Diverse Objects with Simulated HumanoidsZhengyi Luo, Jinkun Cao, Sammy Christen, Alexander Winkler et al.NeurIPS 2024 · 66 citations
- H2O: A Benchmark for Visual Human-human Object Handover AnalysisRuolin Ye, Wenqiang Xu, Zhendong Xue, Tutian Tang et al.ICCV 2021 · 30 citations
- Learning Human-to-Robot Handovers from Point CloudsSammy Joe Christen, Wei Yang, Claudia Pérez-D'Arpino, Otmar Hilliges et al.CVPR 2023
- DemoGrasp: Universal Dexterous Grasping from a Single DemonstrationHaoqi Yuan, Ziye Huang, Ye Wang, Chuan Mao et al.ICLR 2026 · 14 citations
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationNing Gao, Yilun Chen, Shuai Yang, Xinyi Chen et al.CVPR 2025
