DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point Representation
Mohamed Abdelsamad, Michael Ulrich, Bin Yang, Miao Zhang, Yakov Miron, Abhinav Valada
Abstract
Recent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics distribution. In this work, we propose DOS (Distilling Observable Softmaps), a novel SSL framework that self-distills semantic relevance softmaps only at observable (unmasked) points. This strategy prevents information leakage from masked regions and provides richer supervision than discrete token-to-prototype assignments. To address the challenge of unbalanced semantics in an unsupervised setting, we introduce Zipfian prototypes and incorporate them using a modified Sinkhorn-Knopp algorithm, Zipf-Sinkhorn, which enforces a power-law prior over prototype usage and modulates the sharpness of the target softmap during training. DOS outperforms current state-of-the-art methods on semantic segmentation and 3D object detection across multiple benchmarks, including nuScenes, Waymo, SemanticKITTI, ScanNet, and ScanNet200, without relying on extra data or annotations. Our results demonstrate that observable-point softmaps distillation offers a scalable and effective paradigm for learning robust 3D representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e78f5586-d93d-4365-9be3-dd73b42d6bedCited by top-tier papers1
Ask how each one uses itBuilds on12
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Few-Shot Object Detection via Feature ReweightingBingyi Kang, Zhuang Liu, Xin Wang, Fisher Yu et al.ICCV 2019 · 835 citations
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu et al.CVPR 2024 · 31 citations
- GD-MAE: Generative Decoder for MAE Pre-Training on LiDAR Point CloudsHonghui Yang, Tong He, Jiaheng Liu, Hua Chen et al.CVPR 2023
Related papers
- PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point CloudsBarza Nisar, Steven L. WaslanderCVPR 2025
- PointDC: Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel ClusteringZisheng Chen, Hongbin Xu, Weitao Chen, Zhipeng Zhou et al.ICCV 2023 · 21 citations
- Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation LearningRemco F. Leijenaar, Hamidreza KasaeiNeurIPS 2025
- Guided Point Contrastive Learning for Semi-supervised Point Cloud Semantic SegmentationLi Jiang, Shaoshuai Shi, Zhuotao Tian, Xin Lai et al.ICCV 2021 · 137 citations
- Point Cloud Reconstruction Is Insufficient to Learn 3D RepresentationsWeichen Xu, Jian Cao, Tianhao Fu, Ruilong Ren et al.ACM MM 2024 · 1 citation
