Leverage Interactive Affinity for Affordance Learning
Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao, Dacheng Tao
Abstract
Perceiving potential "action possibilities" (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assignment paradigm and presume that there is a unique relationship between functional region and affordance label, yielding poor performance when adapting to unseen environments with large appearance variations. In this paper, we propose to leverage interactive affinity for affordance learning, i.e.extracting interactive affinity from human-object interaction and transferring it to noninteractive objects. Interactive affinity, which represents the contacts between different parts of the human body and local regions of the target object, can provide inherent cues of interconnectivity between humans and objects, thereby reducing the ambiguity of the perceived action possibilities. Specifically, we propose a pose-aided interactive affinity learning framework that exploits human pose to guide the network to learn the interactive affinity from human-object interactions. Particularly, a keypoint heuristic perception (KHP) scheme is devised to exploit the keypoint association of human pose to alleviate the uncertainties due to interaction diversities and contact occlusions. Besides, a contactdriven affordance learning (CAL) dataset is constructed by collecting and labeling over 5, 000 images. Experimental results demonstrate that our method outperforms the representative models regarding objective metrics and visual quality. Code and dataset: github.com/lhc1224/PIAL-Net.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4cdc3fd-2f7d-464b-966d-0ac784716828Cited by top-tier papers6
- LEMON: Learning 3D Human-Object Interaction Relation from 2D ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.CVPR 2024 · 12 citations
- Unlocking 3D Affordance Segmentation with 2D Semantic KnowledgeYu Huang, Zelin Peng, Changsong Wen, Xiaokang Yang et al.CVPR 2026 · 3 citations
- RAGNet: Large-Scale Reasoning-Based Affordance Segmentation Benchmark Towards General GraspingDongming Wu, Yanping Fu, Saike Huang, Yingfei Liu et al.ICCV 2025 · 2 citations
- AffordMatcher: Affordance Learning in 3D Scenes from Visual SignifiersNghia Vu, Tuong Do, Khang Nguyen, Baoru Huang et al.CVPR 2026 · 2 citations
- GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance GroundingYawen Shao, Wei Zhai, Yuhang Yang, Hongchen Luo et al.CVPR 2025
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ViTPose: Simple Vision Transformer Baselines for Human Pose EstimationYufei Xu, Jing Zhang, Qiming Zhang, Dacheng TaoNeurIPS 2022 · 1,105 citations
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive BiasYufei Xu, Qiming Zhang, Jing Zhang, Dacheng TaoNeurIPS 2021 · 429 citations
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
Related papers
- Grounding 3D Object Affordance from 2D Interactions in ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.ICCV 2023 · 69 citations
- AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand PoseJuntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu et al.ICCV 2023 · 78 citations
- Learning Affordance Grounding from Exocentric ImagesHongchen Luo, Wei Zhai, Jing Zhang, Yang Cao et al.CVPR 2022 · 49 citations
- Populating 3D Scenes by Learning Human-Scene InteractionMohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas et al.CVPR 2021
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
