Open-Vocabulary 3D Affordance Understanding via Functional Text Enhancement and Multilevel Representation Alignment
Lin Wu, Wei Wei, Peizhuo Yu, Jianglin Lan
摘要
Understanding 3D affordance is essential for agents to effectively interact with real-world environments, encompassing tasks such as manipulation and navigation. Existing methods typically support open-vocabulary queries through label-based language descriptions but often suffer from limited generalization and weak discriminative ability in their representations. However, affordance understanding requires constructing a coherent semantic landscape from fragmented linguistic expressions-one that preserves intra-class diversity while minimizing inter-class overlap. To address these challenges, we introduce Aff3DFunc, a framework designed to enhance the alignment between affordance and 3D geometry. It begins with a functional text enhancement module grounded in the Information Bottleneck (IB) principle, which strategically enriches affordance semantics by maximizing both relevance and diversity. A dual-encoder architecture is then employed to extract embeddings from both point clouds and text. To bridge the modality gap, we further propose a multilevel representation alignment strategy that incorporates supervised contrastive learning, reinforcing semantic-geometric correspondence in a part-to-whole manner. Extensive experiments demonstrate that our approach significantly enhances the understanding of affordance complexity. The learned representations exhibit high adaptability to diverse text queries, particularly in zero-shot settings. Furthermore, the real-world robot validation confirms that our method improves affordance understanding, enabling more fine-grained manipulation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera 等NeurIPS 2023 · 被引用 371 次
相关 Paper
- QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal EvidenceWeiyu Zhao, Ru Li, Jiaqi Liu, Sizhe Zhao 等CVPR 2026
- PLA: Language-Driven Open-Vocabulary 3D Scene UnderstandingRunyu Ding, Jihan Yang, Chuhui Xue, Wenqing Zhang 等CVPR 2023
- Unlocking 3D Affordance Segmentation with 2D Semantic KnowledgeYu Huang, Zelin Peng, Changsong Wen, Xiaokang Yang 等CVPR 2026 · 被引用 3 次
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu 等NeurIPS 2023 · 被引用 267 次
- SceneForge: Enhancing 3D-text alignment with Structured Scene CompositionsCristian Sbrolli, Matteo MatteucciNeurIPS 2025
