Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
Yu Huang, Zelin Peng, Changsong Wen, Xiaokang Yang, Wei Shen
Abstract
Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recognition. Existing methods, mostly following the paradigm of 3D semantic segmentation or prompt-based frameworks, struggle when geometric cues are weak or ambiguous, as sparse point clouds provide limited functional information. To overcome this limitation, we leverage the rich semantic knowledge embedded in large-scale 2D Vision Foundation Models (VFMs) to guide 3D representation learning through a cross-modal alignment mechanism. Specifically, we propose Cross-Modal Affinity Transfer (CMAT), a pretraining strategy that compels the 3D encoder to align with the semantic structures induced by lifted 2D features. CMAT is driven by a core affinity alignment objective, supported by two auxiliary losses, geometric reconstruction and feature diversity, which together encourage structured and discriminative feature learning. Built upon the CMAT-pretrained backbone, we employ a lightweight affordance segmentor that injects text or visual prompts into the learned 3D space through an efficient cross-attention interface, enabling dense and prompt-aware affordance prediction while preserving the semantic organization established during pretraining. Extensive experiments demonstrate consistent improvements over previous state-of-the-art methods in both accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf666f20-302b-4f4b-9380-cf2b3ff20180Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai et al.NeurIPS 2024 · 179 citations
Related papers
- UniDSeg: Unified Cross-Domain 3D Semantic Segmentation via Visual Foundation Models PriorYao Wu, Mingwei Xing, Yachao Zhang, Xiaotong Luo et al.NeurIPS 2024 · 15 citations
- GEAL: Generalizable 3D Affordance Learning with Cross-Modal ConsistencyDongyue Lu, Lingdong Kong, Tianxin Huang, Gim Hee LeeCVPR 2025
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang et al.ACM MM 2023 · 17 citations
- Point-SAM: Promptable 3D Segmentation Model for Point CloudsYuchen Zhou, Jiayuan Gu, Tung Yen Chiang, Fanbo Xiang et al.ICLR 2025
- EmbodiedSAM: Online Segment Any 3D Thing in Real TimeXiuwei Xu, Huangxing Chen, Linqing Zhao, Ziwei Wang et al.ICLR 2025
