H2O: A Benchmark for Visual Human-human Object Handover Analysis
Ruolin Ye, Wenqiang Xu, Zhendong Xue, Tutian Tang, Yanfeng Wang, Cewu Lu
Abstract
Object handover is a common human collaboration behavior that attracts attention from researchers in Robotics and Cognitive Science. Though visual perception plays an important role in the object handover task, the whole handover process has been specifically explored. In this work, we propose a novel rich-annotated dataset, H2O, for visual analysis of human-human object handovers. The H2O, which contains 18K video clips involving 15 people who hand over 30 objects to each other, is a multi-purpose benchmark. It can support several vision-based tasks, from which, we specifically provide a baseline method, RGPNet, for a less-explored task named Receiver Grasp Prediction. Extensive experiments show that the RGPNet can produce plausible grasps based on the giver's hand-object states in the pre-handover phase. Besides, we also report the hand and object pose errors with existing baselines and show that the dataset can serve as the video demonstrations for robot imitation learning on the handover task. Dataset, model and code will be made public.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96ba8219-a3cc-48c6-bf62-88badefae7d8Cited by top-tier papers12
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu et al.CVPR 2022 · 79 citations
- AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand PoseJuntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu et al.ICCV 2023 · 78 citations
- RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose EstimationLijun Li, Linrui Tian, Xindi Zhang, Qi Wang et al.ICCV 2023 · 28 citations
- HOIDiffusion: Generating Realistic 3D Hand-Object Interaction DataMengqi Zhang, Yang Fu, Zheng Ding, Sifei Liu et al.CVPR 2024 · 18 citations
- TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object UnderstandingYun Liu, Haolin Yang, Xu Si, Ling Liu et al.CVPR 2024 · 12 citations
Builds on2
Related papers
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo et al.ICCV 2021 · 271 citations
- DexH2R: A Benchmark for Dynamic Dexterous Grasping in Human-To-Robot HandoverYouzhuo Wang, Jiayi Ye, Chuyang Xiao, Yiming Zhong et al.ICCV 2025 · 1 citation
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated ManipulationTim Engelbracht, René Zurbrügg, Matteo Wohlrapp, Martin Büchner et al.CVPR 2026 · 9 citations
- GenH2R: Learning Generalizable Human-to-Robot Handover via Scalable Simulation, Demonstration, and ImitationZifan Wang, Junyu Chen, Ziqing Chen, Pengwei Xie et al.CVPR 2024 · 15 citations
- Egocentric Prediction of Action Target in 3DYiming Li, Ziang Cao, Andrew Liang, Benjamin Liang et al.CVPR 2022 · 20 citations
