Hierarchical Intra-Modal Correlation Learning for Label-Free 3D Semantic Segmentation
Xin Kang, Lei Chu, Jiahao Li, Xuejin Chen, Yan Lu
摘要
Recent methods for label-free 3D semantic segmentation aim to assist 3D model training by leveraging the openworld recognition ability of pre-trained vision language models. However, these methods usually suffer from inconsistent and noisy pseudo-labels provided by the vision language models. To address this issue, we present a hierarchical intra-modal correlation learning framework that captures visual and geometric correlations in 3D scenes at three levels: intra-set, intra-scene, and inter-scene, to help learn more compact 3D representations. We refine pseudolabels using intra-set correlations within each geometric consistency set and align features of visually and geometrically similar points using intra-scene and inter-scene correlation learning. We also introduce a feedback mechanism to distill the correlation learning capability into the 3D model. Experiments on both indoor and outdoor datasets show the superiority of our method. We achieve a state-of-the-art 36.6% mIoU on the ScanNet dataset, and a 23.0% mIoU on the nuScenes dataset, with improvements of 7.8% mIoU and 2.2% mIoU compared with previous SOTA. We also provide theoretical analysis and qualitative visualization results to discuss the mechanism and conduct thorough ablation studies to support the effectiveness of our framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AdaCo: Overcoming Visual Foundation Model Noise in 3D Semantic Segmentation via Adaptive Label CorrectionPufan Zou, Shijia Zhao, Weijie Huang, Qiming Xia 等AAAI 2025
- 3D-AVS: LiDAR-based 3D Auto-Vocabulary SegmentationWeijie Wei, Osman Ülger, Fatemeh Karimi Nejadasl, Theo Gevers 等CVPR 2025
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu 等NeurIPS 2022 · 被引用 924 次
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang 等CVPR 2022 · 被引用 494 次
相关 Paper
- Transferring CLIP's Knowledge into Zero-Shot Point Cloud Semantic SegmentationYuanbin Wang, Shaofei Huang, Yulu Gao, Zhen Wang 等ACM MM 2023 · 被引用 17 次
- AGO: Adaptive Grounding for Open World 3D Occupancy PredictionPeizheng Li, Shuxiao Ding, You Zhou, Qingwen Zhang 等ICCV 2025 · 被引用 4 次
- CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIPRunnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu 等CVPR 2023
- RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingJihan Yang, Runyu Ding, Weipeng Deng, Zhe Wang 等CVPR 2024 · 被引用 49 次
- Towards Label-free Scene Understanding by Vision Foundation ModelsRunnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen 等NeurIPS 2023 · 被引用 82 次
