BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
Zhiwei Lin, Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang
摘要
Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive and time-consuming. Self-supervised pre-training is an effective and desirable way to alleviate this dependence on extensive annotated data. In this work, we present BEV-MAE, an efficient masked autoencoder pre-training framework for LiDAR-based 3D object detection in autonomous driving. Specifically, we propose a bird's eye view (BEV) guided masking strategy to guide the 3D encoder learning feature representation in a BEV perspective and avoid complex decoder design during pre-training. Furthermore, we introduce a learnable point token to maintain a consistent receptive field size of the 3D encoder with fine-tuning for masked point cloud inputs. Based on the property of outdoor point clouds in autonomous driving scenarios, i.e., the point clouds of distant objects are more sparse, we propose point density prediction to enable the 3D encoder to learn location information, which is essential for object detection. Experimental results show that BEV-MAE surpasses prior state-of-the-art self-supervised methods and achieves a favorably pre-training efficiency. Furthermore, based on TransFusion-L, BEV-MAE achieves new state-of-the-art LiDAR-based 3D object detection results, with 73.6 NDS and 69.6 mAP on the nuScenes benchmark. The source code will be released at https://github.com/VDIGPKU/BEV-MAE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object DetectionHaoran Zhu, Zhenyuan Dong, Kristi Topollai, Beiyao Sha 等AAAI 2026 · 被引用 5 次
- TurboTrain: Towards Efficient and Balanced Multi-Task Learning for Multi-Agent Perception and PredictionZewei Zhou, Seth Z. Zhao, Tianhui Cai, Zhiyu Huang 等ICCV 2025 · 被引用 2 次
- AnnofreeOD: Detecting All Classes at Low Frame Rates Without Human AnnotationsBoyi Sun, Yuhang Liu, Houxin He, Yonglin Tian 等ICCV 2025 · 被引用 1 次
- Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer UnwrappingZijian Cao, Dayou Zhang, Zeyuan Liu, Zhicheng Liang 等AAAI 2026
它引用的顶会 Paper18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin 等CVPR 2022 · 被引用 1,129 次
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang 等CVPR 2022 · 被引用 794 次
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia 等NeurIPS 2022 · 被引用 762 次
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang 等CVPR 2022 · 被引用 684 次
相关 Paper
- Multi-Scale Neighborhood Occupancy Masked Autoencoder for Self-Supervised Learning in LiDAR Point CloudsMohamed Abdelsamad, Michael Ulrich, Claudius Gläser, Abhinav ValadaCVPR 2025
- DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge DistillationZeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie 等ICCV 2023 · 被引用 65 次
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 被引用 12 次
- A Versatile Multi-View Framework for LiDAR-based 3D Object Detection with Guidance from Panoptic SegmentationHamidreza Fazlali, Yixuan Xu, Yuan Ren, Bingbing LiuCVPR 2022 · 被引用 23 次
- MGTANet: Encoding Sequential LiDAR Points Using Long Short-Term Motion-Guided Temporal Attention for 3D Object DetectionJunho Koh, Junhyung Lee, Youngwoo Lee, Jaekyum Kim 等AAAI 2023 · 被引用 34 次
