2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
Marvin Heidinger, Snehal Jauhri, Vignesh Prasad, Georgia Chalvatzaki
摘要
When interacting with objects, humans effectively reason about which regions of objects are viable for an intended action, i.e., the affordance regions of the object. They can also account for subtle differences in object regions based on the task to be performed and whether one or two hands need to be used. However, current visionbased affordance prediction methods often reduce the problem to naive object part segmentation. In this work, we propose a framework for extracting affordance data from human activity video datasets. Our extracted 2HANDS dataset contains precise object affordance region segmentations and affordance class-labels as narrations of the activity performed. The data also accounts for bimanual actions, i.e., two hands co-ordinating and interacting with one or more objects. We present a VLM-based affordance prediction model, 2HandedAfforder, trained on the dataset and demonstrate superior performance over baselines in affordance region segmentation for various activities. Finally, we show that our predicted affordance regions are actionable, i.e., can be used by an agent performing a task, through demonstration in robotic manipulation scenarios. Project-website: sites.google.com/view/2handedafforder Dataset Image type & source # Images Affordance data Annotation source Annotation type # Aff. classes # Obj. classes Class-labels Bimanual IIT-AFF [34] Exocentric [44] 8.8K Manually-labeled Masks 9 10 Explicit No AGD20K [31] Exo+Egocentric [3, 27] 23.8K Manually-labeled Heatmaps 36 50 Explicit No 3DOI [38] Exo+Egocentric [5, 39] 10K Manually-labeled Points 3 n.a. Explicit No ACP [12] Egocentric [5] 15K Auto-labeled Heatmaps n.a. n.a. none No VRB [1] Egocentric [5] 54K Auto-labeled Heatmaps n.a. n.a. none No 2HANDS Egocentric [5] 278K Auto-labeled Precise Masks 73 163 Narrations Yes
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis 等CVPR 2022 · 被引用 525 次
相关 Paper
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper ManipulationYan Zhao, Ruihai Wu, Zhehuan Chen, Yourong Zhang 等ICLR 2023 · 被引用 2 次
- Understanding Human Hands in Contact at Internet ScaleDandan Shan, Jiaqi Geng, Michelle Shu, David F. FouheyCVPR 2020
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 被引用 26 次
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- Human Hands as Probes for Interactive Object UnderstandingMohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh GuptaCVPR 2022 · 被引用 26 次
