Probing and Bridging Geometry–Interaction Cues for Affordance Reasoning in Vision Foundation Models
Qing Zhang, Xuesong li, Jing Zhang
Abstract
What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zeroshot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action. Our code will be available at: Probing and Bridging Affordance
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1fba9c1-7edd-4db4-aacb-afb569708b77Builds on18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
- Joint Hand Motion and Interaction Hotspots Prediction from Egocentric VideosShaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong WangCVPR 2022 · 69 citations
Related papers
- Unlocking 3D Affordance Segmentation with 2D Semantic KnowledgeYu Huang, Zelin Peng, Changsong Wen, Xiaokang Yang et al.CVPR 2026 · 3 citations
- Do Computer Vision Foundation Models Learn the Low-level Characteristics of the Human Visual System?Yancheng Cai, Fei Yin, Dounia Hammou, Rafal MantiukCVPR 2025
- Feat2GS: Probing Visual Foundation Models with Gaussian SplattingYue Chen, Xingyu Chen, Anpei Chen, Gerard Pons-Moll et al.CVPR 2025
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic ManipulationZhihao Zhu, Yifan Zheng, Siyu Pan, Yaohui Jin et al.ICCV 2025
