Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
Xumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang, Jie Zhou, Jiwen Lu
Abstract
We present Point-BERT, a new paradigm for learning Transformers to generalize the concept of BERT [8] to 3D point cloud. Inspired by BERT, we devise a Masked Point Modeling (MPM) task to pre-train point cloud Transformers. Specifically, we first divide a point cloud into several local point patches, and a point cloud Tokenizer with a discrete Variational AutoEncoder (dVAE) is designed to generate discrete point tokens containing meaningful local information. Then, we randomly mask out some patches of input point clouds and feed them into the backbone Transformers. The pre-training objective is to recover the original point tokens at the masked locations under the supervision of point tokens obtained by the Tokenizer. Extensive experiments demonstrate that the proposed BERT-style pre-training strategy significantly improves the performance of standard point cloud Transformers. Equipped with our pre-training strategy, we show that a pure Transformer architecture attains 93.8% accuracy on ModelNet40 and 83.1% accuracy on the hardest setting of ScanObjectNN, surpassing carefully designed point cloud models with much fewer hand-made designs. We also demonstrate that the representations learned by Point-BERT transfer well to new tasks and domains, where our models largely advance the state-of-the-art of few-shot point cloud classification task. The code and pre-trained models are available at https://github.com/lulutang0608/Point-BERT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0dc8ffb-ffd7-4d30-b312-35684af8359aCited by top-tier papers287
- PointNeXt: Revisiting PointNet++ with Improved Training and Scaling StrategiesGuocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai et al.NeurIPS 2022 · 1,270 citations
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu et al.NeurIPS 2024 · 380 citations
- An Embodied Generalist Agent in 3D WorldJiangyong Huang, Silong Yong, Xiaojian Ma, Xiongkun Linghu et al.ICML 2024 · 361 citations
- OpenShape: Scaling Up 3D Shape Representation Towards Open-World UnderstandingMinghua Liu, Ruoxi Shi, Kaiming Kuang, Yinhao Zhu et al.NeurIPS 2023 · 267 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- General Point Model Pretraining with Autoencoding and AutoregressiveZhe Li, Zhangyang Gao, Cheng Tan, Bocheng Ren et al.CVPR 2024
- PointGPT: Auto-regressively Generative Pre-training from Point CloudsGuangyan Chen, Meiling Wang, Yi Yang, Kai Yu et al.NeurIPS 2023 · 219 citations
- Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked ModelingKeyu Tian, Yi Jiang, Qishuai Diao, Chen Lin et al.ICLR 2023 · 37 citations
- Frozen CLIP Transformer Is an Efficient Point Cloud EncoderXiaoshui Huang, Zhou Huang, Sheng Li, Wentao Qu et al.AAAI 2024 · 30 citations
- Take-A-Photo: 3D-to-2D Generative Pre-training of Point Cloud ModelsZiyi Wang, Xumin Yu, Yongming Rao, Jie Zhou et al.ICCV 2023 · 34 citations
