Geometrically-Driven Aggregation for Zero-Shot 3D Point Cloud Understanding
Guofeng Mei, Luigi Riz, Yiming Wang, Fabio Poiesi
摘要
Zero-shot 3D point cloud understanding can be achieved via 2D Vision-Language Models (VLMs). Existing strategies directly map VLM representations from 2D pixels of rendered or captured views to 3D points, overlooking the inherent and expressible point cloud geometric structure. Geometrically similar or close regions can be exploited for bolstering point cloud understanding as they are likely to share semantic information. To this end, we introduce the first training-free aggregation technique that leverages the point cloud's 3D geometric structure to improve the quality of the transferred VLM representations. Our approach operates iteratively, performing local-to-global aggregation based on geometric and semantic point-level reasoning. We benchmark our approach on three downstream tasks, including classification, part segmentation, and semantic segmentation, with a variety of datasets representing both synthetic/real-world, and indoor/outdoor scenarios. Our approach achieves new state-of-the-art results in all benchmarks. Code and dataset are available at https: //luigiriz.github.io/geoze-website/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic SegmentationXujing Tao, Chuxin Wang, Yubo Ai, Zhixin Cheng 等CVPR 2026 · 被引用 3 次
- Real-time 3D Object Detection with Inference-Aligned LearningChenyu Zhao, Xianwei Zheng, Zimin Xia, Linwei Yue 等AAAI 2026 · 被引用 1 次
- SegGraph: Leveraging Graphs of SAM Segments for Few-Shot 3D Part SegmentationYueyang Hu, Haiyong Jiang, Haoxuan Song, Jun Xiao 等NeurIPS 2025 · 被引用 1 次
- Multimodality Helps Few-shot 3D Point Cloud Semantic SegmentationZhaochong An, Guolei Sun, Yun Liu, Runjia Li 等ICLR 2025
- Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score PropagationTiankai Chen, Yushu Li, Adam Goodge, Fei Teng 等ICCV 2025
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen 等ICCV 2019 · 被引用 1,003 次
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu 等NeurIPS 2022 · 被引用 924 次
- Rethinking Network Design and Local Geometry in Point Cloud: A Simple Residual MLP FrameworkXu Ma, Can Qin, Haoxuan You, Haoxi Ran 等ICLR 2022 · 被引用 841 次
相关 Paper
- Point2Real: Bridging the Gap between Point Cloud and Realistic Image for Open-World 3D RecognitionHanxuan Li, Bin Fu, Ruiping Wang, Xilin ChenAAAI 2024 · 被引用 1 次
- Zero-Shot 3D Question Answering via Hierarchical View-to-Token TransportationDongsheng Wang, Dawei Su, Hui HuangICML 2026
- Bridging Language and Geometric Primitives for Zero-shot Point Cloud SegmentationRunnan Chen, Xinge Zhu, Nenglun Chen, Wei Li 等ACM MM 2023 · 被引用 8 次
- UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language ConditionsWenbin Tan, Jiawen Lin, Yuan Xie, Yachao Zhang 等CVPR 2026
- Affinity3D: Propagating Instance-Level Semantic Affinity for Zero-Shot Point Cloud Semantic SegmentationHaizhuang Liu, Junbao Zhuo, Chen Liang, Jiansheng Chen 等ACM MM 2024 · 被引用 2 次
