Geometrically-Driven Aggregation for Zero-Shot 3D Point Cloud Understanding
Guofeng Mei, Luigi Riz, Yiming Wang, Fabio Poiesi
Abstract
Zero-shot 3D point cloud understanding can be achieved via 2D Vision-Language Models (VLMs). Existing strategies directly map VLM representations from 2D pixels of rendered or captured views to 3D points, overlooking the inherent and expressible point cloud geometric structure. Geometrically similar or close regions can be exploited for bolstering point cloud understanding as they are likely to share semantic information. To this end, we introduce the first training-free aggregation technique that leverages the point cloud's 3D geometric structure to improve the quality of the transferred VLM representations. Our approach operates iteratively, performing local-to-global aggregation based on geometric and semantic point-level reasoning. We benchmark our approach on three downstream tasks, including classification, part segmentation, and semantic segmentation, with a variety of datasets representing both synthetic/real-world, and indoor/outdoor scenarios. Our approach achieves new state-of-the-art results in all benchmarks. Code and dataset are available at https: //luigiriz.github.io/geoze-website/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0aa52b9-e200-4d6f-9d64-3d0ca6e561e4Cited by top-tier papers6
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng et al.CVPR 2026 · 3 citations
- Real-time 3D Object Detection with Inference-Aligned LearningChenyu Zhao, Xianwei Zheng, Zimin Xia, Linwei Yue et al.AAAI 2026 · 1 citation
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part SegmentationYueyang Hu, Haiyong Jiang, Haoxuan Song, Jun Xiao et al.NeurIPS 2025 · 1 citation
- Multimodality Helps Few-shot 3D Point Cloud Semantic SegmentationZhaochong An, Guolei Sun, Yun Liu, Runjia Li et al.ICLR 2025
- Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationTiankai Chen, Yushu Li, Adam Goodge, Fei Teng et al.ICCV 2025
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP FrameworkXu Ma, Can Qin, Haoxuan You, Haoxi Ran et al.ICLR 2022 · 841 citations
Related papers
- Point2Real: Bridging the Gap between Point Cloud and Realistic Image for Open-World 3D RecognitionHanxuan Li, Bin Fu, Ruiping Wang, Xilin ChenAAAI 2024 · 1 citation
- Zero-Shot 3D Question Answering via Hierarchical View-to-Token TransportationDongsheng Wang, Dawei Su, Hui HuangICML 2026
- Bridging Language and Geometric Primitives for Zero-shot Point Cloud SegmentationRunnan Chen, Xinge Zhu, Nenglun Chen, Wei Li et al.ACM MM 2023 · 8 citations
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang et al.CVPR 2026
- Affinity3D: Propagating Instance-Level Semantic Affinity for Zero-Shot Point Cloud Semantic SegmentationHaizhuang Liu, Junbao Zhuo, Chen Liang, Jiansheng Chen et al.ACM MM 2024 · 2 citations
