Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
Bin Yang, Mohamed Abdelsamad, Miao Zhang, Alexandru Paul Condurache
摘要
Recent advances in self-supervised learning (SSL) for point clouds have substantially improved 3D scene understanding without human annotations. Existing approaches emphasize semantic awareness by enforcing feature consistency across augmented views or by masked scene modeling. However, the resulting representations transfer poorly to localization, and often require full finetuning for strong performance. Accurate localization is a fundamental component of 3D perception, thus bridging this gap is crucial for progressing toward true 3D foundation models that supports all downstream tasks on 3D data. In this work, we introduce PointINS, a localization-oriented self-supervised framework that enriches point cloud representations through geometry-aware learning. PointINS employs an orthogonal localization branch to jointly learn high-level semantic understanding and geometric reasoning, yielding localization awareness. We identify two consistent properties essential for robust localization and formulate them as complementary regularization strategies, Offset Distribution Regularization (ODR), which aligns predicted offsets with empirically observed geometric priors, and Spatial Clustering Regularization (SCR), which enforces local coherence by regularizing offsets with pseudo instance masks. Through extensive experiments across five datasets, PointINS achieves on average +3.5% mAP improvement for indoor instance segmentation and +4.1% PQ gain for outdoor panoptic segmentation, paving the way for scalable 3D foundation models. Code will be released upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper28
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 被引用 333 次
- SoftGroup for 3D Instance Segmentation on Point CloudsThang Vu, Kookhoi Kim, Tung Minh Luu, Thanh Xuan Nguyen 等CVPR 2022 · 被引用 251 次
相关 Paper
- PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point CloudsBarza Nisar, Steven L. WaslanderCVPR 2025
- PointClustering: Unsupervised Point Cloud Pre-training using Transformation Invariance in ClusteringFuchen Long, Ting Yao, Zhaofan Qiu, Lusong Li 等CVPR 2023
- Clustering based Point Cloud Representation Learning for 3D AnalysisTuo Feng, Wenguan Wang, Xiaohan Wang, Yi Yang 等ICCV 2023 · 被引用 53 次
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 被引用 12 次
- Point-SAM: Promptable 3D Segmentation Model for Point CloudsYuchen Zhou, Jiayuan Gu, Tung Yen Chiang, Fanbo Xiang 等ICLR 2025
