Video-to-Image Affordance Grounding via Visual Conceptual Learning
Zhiyuan Fan, Keyi Liang
Abstract
Video-to-Image Affordance Grounding aims to localize object affordances in static images by learning from human demonstration videos. However, existing fully supervised methods rely on paired image-video inputs during both training and inference, which significantly limits their practicality in real-world scenarios. Conversely, weakly supervised approaches support video-free inference but often struggle to capture the most critical interaction features necessary for affordance learning. To overcome these limitations, we propose a novel two-stage framework, VCL. In the first stage, we adopt the standard fully supervised paradigm to train the model, enabling it to effectively extract meaningful interaction features from demonstration videos. In the second stage, these extracted features are conceptualized, and a conceptual module is introduced to map natural language instructions into the concept space. Our approach establishes a new paradigm for affordance learning from demonstrations, enabling the model to learn from video but perform precise text-conditioned inference. Experiments on two benchmark datasets show that our base model outperforms previous state-of-the-art methods, while our text-conditioned model achieves competitive performance without requiring paired video inputs during inference. Code will be released at https://github.com/Fanzy27/VCL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 218fa283-1e7a-48ed-975b-5c055b9967bbRelated papers
- Grounded Human-Object Interaction Hotspots From VideoTushar Nagarajan, Christoph Feichtenhofer, Kristen GraumanICCV 2019 · 194 citations
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingWonJun Moon, Hyun Seok Seong, Jae-Pil HeoICCV 2025 · 1 citation
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic PriorsPeiran Xu, Yadong MuICLR 2025
- Interventional Video Grounding With Dual Contrastive LearningGuoshun Nan, Rui Qiao, Yao Xiao, Jun Liu et al.CVPR 2021
- Weakly Supervised Multimodal Affordance Grounding for Egocentric ImagesLingjing Xu, Yang Gao, Wenfeng Song, Aimin HaoAAAI 2024 · 19 citations
