Intermediate Connectors and Geometric Priors for Language-Guided Affordance Segmentation on Unseen Object Categories
Yicong Li, Yiyang Chen, Zhenyuan Ma, Junbin Xiao, Xiang Wang, Angela Yao
Abstract
Language-guided Affordance Segmentation (LASO) aims to identify actionable object regions based on text instructions. At the core of its practicality is learning generalizable affordance knowledge that captures functional regions across diverse objects. However, current LASO solutions struggle to extend learned affordances to object categories that are not encountered during training. Scrutinizing these designs, we identify limited generalizability on unseen categories, stemming from (1) underutilized generalizable patterns in the intermediate layers of both 3D and text backbones, which impedes the formation of robust affordance knowledge, and (2) the inability to handle substantial variability in affordance regions across object categories due to a lack of structural knowledge of the target region. Towards this, we introduce a GeneraLized frAmework on uNseen CategoriEs (GLANCE), incorporating two key components: a cross-modal connector that links intermediate stages of the text and 3D backbones to enrich pointwise embeddings with affordance concepts, and a VLM-guided query generator that provides affordance priors by extracting a few 3D key points based on the intra-view reliability and crossview consistency of their multi-view segmentation masks. Extensive experiments on two benchmark datasets demonstrate that GLANCE outperforms state-of-the-art methods (SoTAs), with notable improvements in generalization to unseen categories. Our code is available at https:// github.com/Monoxide-Chen/Affordance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddf8b024-ba43-4c17-b425-f972cbbb5629Cited by top-tier papers2
- Geometric Alignment and Prior Modulation for View-Guided Point Cloud Completion on Unseen CategoriesJingqiao Xiu, Yicong Li, Na Zhao, Han Fang et al.ICCV 2025 · 2 citations
- RelaxFlow: Text-Driven Amodal 3D GenerationJiayin Zhu, Guoji Fu, Xiaolu Liu, Qiyuan He et al.ICML 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- Video as Conditional Graph Hierarchy for Multi-Granular Question AnsweringJunbin Xiao, Angela Yao, Zhiyuan Liu, Yicong Li et al.AAAI 2022 · 145 citations
Related papers
- LASO: Language-Guided Affordance Segmentation on 3D ObjectYicong Li, Na Zhao, Junbin Xiao, Chun Feng et al.CVPR 2024
- ViSPLA: Visual Iterative Self-Prompting for Language-Guided 3D Affordance LearningHritam Basak, Zhaozheng YinNeurIPS 2025 · 2 citations
- 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D WorldsHengshuo Chu, Xiang Deng, Qi Lv, Xiaoyang Chen et al.ICLR 2025
- GEAL: Generalizable 3D Affordance Learning with Cross-Modal ConsistencyDongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee LeeCVPR 2025
- MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsDongpan Chen, Dehui Kong, Jinghua Li, Baocai YinAAAI 2025 · 5 citations
