GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training
Xiaoyu Tian, Haoxi Ran, Yue Wang, Hang Zhao
Abstract
This paper tries to address a fundamental question in point cloud self-supervised learning: what is a good signal we should leverage to learn features from point clouds without annotations? To answer that, we introduce a point cloud representation learning framework, based on geometric feature reconstruction. In contrast to recent papers that directly adopt masked autoencoder (MAE) and only predict original coordinates or occupancy from masked point clouds, our method revisits differences between images and point clouds and identifies three self-supervised learning objectives peculiar to point clouds, namely centroid prediction, normal estimation, and curvature prediction. Combined with occupancy prediction, these four objectives yield a nontrivial self-supervised learning task and mutually facilitate models to better reason fine-grained geometry of point clouds. Our pipeline is conceptually simple and it consists of two major steps: first, it randomly masks out groups of points, followed by a Transformer-based point cloud encoder; second, a lightweight Transformer decoder predicts centroid, normal, and curvature for points in each voxel. We transfer the pre-trained Transformer encoder to a downstream peception model. On the nuScene Datset, our model achieves 3.38 mAP improvement for object detection, 2.1 mIoU gain for segmentation, and 1.7 AMOTA gain for multi-object tracking. We also conduct experiments on the Waymo Open Dataset and achieve significant performance improvements over baselines as well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Towards Compact 3D Representations via Point Feature Enhancement Masked AutoencodersYaohua Zha, Huizhen Ji, Jinmin Li, Rongsheng Li et al.AAAI 2024 · 66 citations
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu et al.CVPR 2024 · 31 citations
- Annotator: A Generic Active Learning Baseline for LiDAR Semantic SegmentationBinhui Xie, Shuang Li, Qingju Guo, Chi Harold Liu et al.NeurIPS 2023 · 21 citations
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingChen Min, Dawei Zhao, Liang Xiao, Jian Zhao et al.CVPR 2024 · 20 citations
- Pre-training with Random Orthogonal Projection Image ModelingMaryam Haghighat, Peyman Moghadam, Shaheer Mohamed, Piotr KoniuszICLR 2024 · 15 citations
Builds on27
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
Related papers
- PCP-MAE: Learning to Predict Centers for Point Masked AutoencodersXiangdong Zhang, Shaofeng Zhang, Junchi YanNeurIPS 2024 · 44 citations
- Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point CloudsMohamed Abdelsamad, Michael Ulrich, Claudius Gläser, Abhinav ValadaCVPR 2025
- 3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud PretrainingSiming Yan, Yuqi Yang, Yu-Xiao Guo, Hao Pan et al.ICLR 2024 · 21 citations
- Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-trainingXiaoyang Xiao, Runzhao Yao, Zhiqiang Tian, Shaoyi DuNeurIPS 2025 · 4 citations
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
