Weakly Supervised Multimodal Affordance Grounding for Egocentric Images
Lingjing Xu, Yang Gao, Wenfeng Song, Aimin Hao
摘要
To enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Interaction-aware Representation Modeling With Co-Occurrence Consistency for Egocentric Hand-Object ParsingYUEJIAO SU, Yi Wang, Lei Yao, Yawen Cui 等ICLR 2026 · 被引用 5 次
- Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric RefinementLian He, Meng Liu, Qilang Ye, Yu Zhou 等AAAI 2026 · 被引用 3 次
- OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part DetectionHeng Su, Mengying Xie, Nieqing Cao, Yan Ding 等ICCV 2025 · 被引用 2 次
- Closed-Loop Transfer for Weakly-Supervised Affordance GroundingJiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei YangICCV 2025 · 被引用 1 次
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingWonJun Moon, Hyun Seok Seong, Jae-Pil HeoICCV 2025 · 被引用 1 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang 等CVPR 2022 · 被引用 527 次
相关 Paper
- Learning Affordance Grounding from Exocentric ImagesHongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 等CVPR 2022 · 被引用 49 次
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
- Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingYuxuan Wang, Aming Wu, Muli Yang, Yukuan Min 等CVPR 2025
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 被引用 194 次
- HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance GroundingLei Yao, Yong Chen, Yuejiao Su, Yi Wang 等CVPR 2026
