2HandedAfforder: Learning Precise Actionable Bimanual Affordances from Human Videos
Marvin Heidinger, Snehal Jauhri, Vignesh Prasad, Georgia Chalvatzaki
Abstract
When interacting with objects, humans effectively reason about which regions of objects are viable for an intended action, i.e., the affordance regions of the object. They can also account for subtle differences in object regions based on the task to be performed and whether one or two hands need to be used. However, current visionbased affordance prediction methods often reduce the problem to naive object part segmentation. In this work, we propose a framework for extracting affordance data from human activity video datasets. Our extracted 2HANDS dataset contains precise object affordance region segmentations and affordance class-labels as narrations of the activity performed. The data also accounts for bimanual actions, i.e., two hands co-ordinating and interacting with one or more objects. We present a VLM-based affordance prediction model, 2HandedAfforder, trained on the dataset and demonstrate superior performance over baselines in affordance region segmentation for various activities. Finally, we show that our predicted affordance regions are actionable, i.e., can be used by an agent performing a task, through demonstration in robotic manipulation scenarios. Project-website: sites.google.com/view/2handedafforder Dataset Image type & source # Images Affordance data Annotation source Annotation type # Aff. classes # Obj. classes Class-labels Bimanual IIT-AFF [34] Exocentric [44] 8.8K Manually-labeled Masks 9 10 Explicit No AGD20K [31] Exo+Egocentric [3, 27] 23.8K Manually-labeled Heatmaps 36 50 Explicit No 3DOI [38] Exo+Egocentric [5, 39] 10K Manually-labeled Points 3 n.a. Explicit No ACP [12] Egocentric [5] 15K Auto-labeled Heatmaps n.a. n.a. none No VRB [1] Egocentric [5] 54K Auto-labeled Heatmaps n.a. n.a. none No 2HANDS Egocentric [5] 278K Auto-labeled Precise Masks 73 163 Narrations Yes
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4ff3ca2-fbab-403b-8a2c-75323abb4eacCited by top-tier papers1
Ask how each one uses itBuilds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
Related papers
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper ManipulationYan Zhao, Ruihai Wu, Zhehuan Chen, Yourong Zhang et al.ICLR 2023 · 2 citations
- Understanding Human Hands in Contact at Internet ScaleDandan Shan, Jiaqi Geng, Michelle Shu, David F. FouheyCVPR 2020
- Multi-label affordance mapping from egocentric visionLorenzo Mur-Labadia, Josechu J. Guerrero, Ruben Martinez-CantinICCV 2023 · 26 citations
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
- Human Hands as Probes for Interactive Object UnderstandingMohit Goyal, Sahil Modi, Rishabh Goyal, Saurabh GuptaCVPR 2022 · 26 citations
