Weakly Supervised Multimodal Affordance Grounding for Egocentric Images
Lingjing Xu, Yang Gao, Wenfeng Song, Aimin Hao
Abstract
To enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ec2cc7f-a18b-425a-9bea-bf0a5839e775Cited by top-tier papers9
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingYUEJIAO SU, Yi Wang, Lei Yao, Yawen Cui et al.ICLR 2026 · 5 citations
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric RefinementLian He, Meng Liu, Qilang Ye, Yu Zhou et al.AAAI 2026 · 3 citations
- OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part DetectionHeng Su, Mengying Xie, Nieqing Cao, Yan Ding et al.ICCV 2025 · 2 citations
- Closed-Loop Transfer for Weakly-Supervised Affordance GroundingJiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei YangICCV 2025 · 1 citation
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingWonJun Moon, Hyun Seok Seong, Jae-Pil HeoICCV 2025 · 1 citation
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
Related papers
- Learning Affordance Grounding from Exocentric ImagesHongchen Luo, Wei Zhai, Jing Zhang, Yang Cao et al.CVPR 2022 · 49 citations
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
- Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingYuxuan Wang, Aming Wu, Muli Yang, Yukuan Min et al.CVPR 2025
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
- HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance GroundingLei Yao, Yong Chen, Yuejiao Su, Yi Wang et al.CVPR 2026
