Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
Tim Engelbracht, René Zurbrügg, Matteo Wohlrapp, Martin Büchner, Abhinav Valada, Marc Pollefeys, Hermann Blum, Zuria Bauer
Abstract
We present a dataset for force-grounded, cross-view articulated manipulation that couples what is seen with what is done and what is felt during real human interaction. The dataset contains 3048 sequences across 381 articulated objects in 38 environments. Each object is operated under four embodiments - (i) human hand, (ii) human hand with a wrist-mounted camera, (iii) handheld UMI gripper, and (iv) a custom Hoi! gripper - where the tool embodiment provide synchronized end-effector forces and tactile sensing. Our dataset offers a holistic view of interaction understanding from video, enabling researchers to evaluate how well methods transfer between human and robotic viewpoints, but also investigate underexplored modalities such as force sensing and prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32a330d3-9bdb-4114-9368-7d088a0a9376Builds on11
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa et al.CVPR 2024 · 110 citations
- Ditto: Building Digital Twins of Articulated Objects from InteractionZhenyu Jiang, Cheng-Chun Hsu, Yuke ZhuCVPR 2022 · 77 citations
- AKB-48: A Real-World Articulated Object Knowledge BaseLiu Liu, Wenqiang Xu, Haoyuan Fu, Sucheng Qian et al.CVPR 2022 · 64 citations
Related papers
- Learning Object-Centric Motion Priors from Human for Robotic Dexterous ManipulationZhengdong Hong, Guofeng ZhangAAAI 2026
- Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing GlovesXinyu Zhang, Ziyi Kou, Chuan Qin, Mia Huang et al.CVPR 2026 · 5 citations
- H2O: A Benchmark for Visual Human-human Object Handover AnalysisRuolin Ye, Wenqiang Xu, Zhendong Xue, Tutian Tang et al.ICCV 2021 · 30 citations
- RoboWheel: A Data Engine from Real-World Human Demonstrations for Cross-Embodiment Robotic LearningYuhong Zhang, Zihan Gao, Shengpeng Li, Ling-Hao Chen et al.CVPR 2026 · 11 citations
- I'M HOI: Inertia-Aware Monocular Capture of 3D Human-Object InteractionsChengfeng Zhao, Juze Zhang, Jiashen Du, Ziwei Shan et al.CVPR 2024 · 9 citations
