Human Hands as Probes for Interactive Object Understanding
Mohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh Gupta
摘要
Interactive object understanding, or what we can do to objects and how is a long-standing goal of computer vision. In this paper, we tackle this problem through observation of human hands in in-the-wild egocentric videos. We demonstrate that observation of what human hands interact with and how can provide both the relevant data and the necessary supervision. Attending to hands, readily localizes and stabilizes active objects for learning and reveals places where interactions with objects occur. Analyzing the hands shows what we can do to objects and how. We apply these basic principles on the EPIC-KITCHENS dataset, and successfully learn state-sensitive features, and object affordances (regions of interaction and afforded grasps), purely by observing hands in egocentric videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
- Towards A Richer 2D Understanding of Hands at ScaleTianyi Cheng, Dandan Shan, Ayda Hassen, Richard E. L. Higgins 等NeurIPS 2023 · 被引用 48 次
- Pretrained Language Models as Visual Planners for Human AssistanceDhruvesh Patel, Hamid Eghbalzadeh, Nitin Kamra, Michael Louis Iuzzolino 等ICCV 2023 · 被引用 41 次
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang 等NeurIPS 2024 · 被引用 38 次
- Understanding 3D Object Interaction from a Single ImageShengyi Qian, David F. FouheyICCV 2023 · 被引用 35 次
它引用的顶会 Paper19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 被引用 673 次
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo 等ICCV 2021 · 被引用 271 次
相关 Paper
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- Shaping embodied agent behavior with activity-context priors from egocentric videoTushar Nagarajan, Kristen GraumanNeurIPS 2021 · 被引用 23 次
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 被引用 26 次
- Understanding Human Hands in Contact at Internet ScaleDandan Shan, Jiaqi Geng, Michelle Shu, David F. FouheyCVPR 2020
- Opening the Vocabulary of Egocentric ActionsDibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela YaoNeurIPS 2023 · 被引用 28 次
