Understanding Human Hands in Contact at Internet Scale
Dandan Shan, Jiaqi Geng, Michelle Shu, David F. Fouhey
摘要
Hands are the central means by which humans manipulate their world and being able to reliably extract hand state information from Internet videos of humans engaged in their hands has the potential to pave the way to systems that can learn from petabytes of video data. This paper proposes steps towards this by inferring a rich representation of hands engaged in interaction method that includes: hand location, side, contact state, and a box around the object in contact. To support this effort, we gather a large-scale dataset of hands in contact with objects consisting of 131 days of footage as well as a 100K annotated hand-contact video frame dataset. The learned model on this dataset can serve as a foundation for handcontact understanding in videos. We quantitatively evaluate it both on its own and in service of predicting and learning from 3D meshes of human hands.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper120
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 等NeurIPS 2023 · 被引用 336 次
- Reconstructing Hand-Object Interactions in the WildZhe Cao, Ilija Radosavovic, Angjoo Kanazawa, Jitendra MalikICCV 2021 · 被引用 184 次
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D ConsistencyDevendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta, Jitendra Malik 等NeurIPS 2021 · 被引用 100 次
- PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingJang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras 等NeurIPS 2025 · 被引用 97 次
它引用的顶会 Paper4
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell 等ICCV 2019 · 被引用 493 次
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- Contextual Attention for Hand Detection in the WildSupreeth Narasimhaswamy, Zhengwei Wei, Yang Wang, Justin Zhang 等ICCV 2019 · 被引用 63 次
相关 Paper
- Towards A Richer 2D Understanding of Hands at ScaleTianyi Cheng, Dandan Shan, Ayda Hassen, Richard E. L. Higgins 等NeurIPS 2023 · 被引用 48 次
- Human Hands as Probes for Interactive Object UnderstandingMohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh GuptaCVPR 2022 · 被引用 26 次
- HOnnotate: A Method for 3D Annotation of Hand and Object PosesShreyas Hampali, Mahdi Rad, Markus Oberweger, Vincent LepetitCVPR 2020
- InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction GenerationSirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu 等CVPR 2025
- Detecting Hands and Recognizing Physical Contact in the WildSupreeth Narasimhaswamy, Trung Nguyen, Minh Hoai NguyenNeurIPS 2020 · 被引用 57 次
