Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance Grounding
Yuxuan Wang, Aming Wu, Muli Yang, Yukuan Min, Yihang Zhu, Cheng Deng
Abstract
This paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixellevel annotations. Most existing methods usually consider the affordance regions to be isolated and directly employ class activation maps to conduct localization, ignoring the relationships with other object components and weakening the performance. For example, a cup's handle is combined with its body to achieve the pouring ability. Obviously, capturing the region relationships is beneficial for improving the localization accuracy of affordance regions. To this end, we first explore exploiting hypergraph to discover these relations and propose a Reasoning Mamba (R-Mamba) framework. We first extract feature embeddings from exocentric and egocentric images to construct the hypergraphs consisting of multiple vertices and hyperedges, which capture the in-context local region relationships between different visual components. Subsequently, we design a Hypergraph-guided State Space (HSS) block to reorganize these local relationships from the global perspective. By this mechanism, the model could leverage the captured relationships to improve the localization accuracy of affordance regions. Extensive experiments and visualization analyses demonstrate the superiority of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e9d0875-5390-4abc-adfc-387c9b15de2eCited by top-tier papers4
- VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual GroundingYihang Zhu, Jinhao Zhang, Yuxuan Wang, Aming Wu et al.ICCV 2025 · 1 citation
- Entangled No More: Multi-Domain Decoupling for Robust Dynamic Graph Neural NetworksYouda Mo, Chaobo He, Junwei Cheng, Peng Mei et al.ICML 2026
- Learning Attribute–Affordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance GroundingYuxuan Wang, Tong Li, Yihang Zhu, Guangtao Lyu et al.ICML 2026
- Probing and Bridging Geometry–Interaction Cues for Affordance Reasoning in Vision Foundation ModelsQing Zhang, Xuesong li, Jing ZhangCVPR 2026
Builds on16
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
Related papers
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
- Weakly Supervised Multimodal Affordance Grounding for Egocentric ImagesLingjing Xu, Yang Gao, Wenfeng Song, Aimin HaoAAAI 2024 · 19 citations
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingWonJun Moon, Hyun Seok Seong, Jae-Pil HeoICCV 2025 · 1 citation
- LOCATE: Localize and Transfer Object Parts for Weakly Supervised Affordance GroundingGen Li, Varun Jampani, Deqing Sun, Laura Sevilla-LaraCVPR 2023
- Learning Affordance Grounding from Exocentric ImagesHongchen Luo, Wei Zhai, Jing Zhang, Yang Cao et al.CVPR 2022 · 49 citations
