Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point Clouds
Mohamed Abdelsamad, Michael Ulrich, Claudius Gläser, Abhinav Valada
Abstract
Masked autoencoders (MAE) have shown tremendous potential for self-supervised learning (SSL) in vision and beyond. However, point clouds from LiDARs used in automated driving are particularly challenging for MAEs since large areas of the 3D volume are empty. Consequently, existing work suffers from leaking occupancy information into the decoder and has significant computational complexity, thereby limiting the SSL pre-training to only 2D bird's eye view encoders in practice. In this work, we propose the novel neighborhood occupancy MAE (NOMAE) that overcomes the aforementioned challenges by employing masked occupancy reconstruction only in the neighborhood of nonmasked voxels. We incorporate voxel masking and occupancy reconstruction at multiple scales with our proposed hierarchical mask generation technique to capture features of objects of different sizes in the point cloud. NOMAEs are extremely flexible and can be directly employed for SSL in existing 3D architectures. We perform extensive evaluations on the nuScenes and Waymo Open datasets for the downstream perception tasks of semantic segmentation and 3D object detection, comparing with both discriminative and generative SSL methods. The results demonstrate that NO-MAE sets the new state-of-the-art on multiple benchmarks for multiple point cloud perception tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69eefd8e-be48-4b49-b6fd-a55e28fdbe71Cited by top-tier papers5
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object DetectionHaoran Zhu, Zhenyuan Dong, Kristi Topollai, Beiyao Sha et al.AAAI 2026 · 5 citations
- Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point CloudsBin Yang, Mohamed Abdelsamad, Miao Zhang, Alexandru Paul ConduracheCVPR 2026 · 4 citations
- Dynamic Focused Masking for Autoregressive Embodied Occupancy PredictionYuan Sun, Julio Contreras, Jorge OrtizNeurIPS 2025 · 3 citations
- Collaborative Learning for Semi-Supervised LiDAR Semantic SegmentationBin Yang, Alexandru Paul ConduracheICML 2026 · 1 citation
- DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point RepresentationMohamed Abdelsamad, Michael Ulrich, Bin Yang, Miao Zhang et al.AAAI 2026 · 1 citation
Builds on28
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
Related papers
- BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving ScenariosZhiwei Lin, Yongtao Wang, Shengxiang Qi, Nan Dong et al.AAAI 2024 · 32 citations
- GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-TrainingXiaoyu Tian, Haoxi Ran, Yue Wang, Hang ZhaoCVPR 2023
- 3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud PretrainingSiming Yan, Yuqi Yang, Yu-Xiao Guo, Hao Pan et al.ICLR 2024 · 21 citations
- Point Cloud Reconstruction Is Insufficient to Learn 3D RepresentationsWeichen Xu, Jian Cao, Tianhao Fu, Ruilong Ren et al.ACM MM 2024 · 1 citation
- PSA-SSL: Pose and Size-aware Self-Supervised Learning on LiDAR Point CloudsBarza Nisar, Steven L. WaslanderCVPR 2025
