MV-JAR: Masked Voxel Jigsaw and Reconstruction for LiDAR-Based Self-Supervised Pre-Training
Runsen Xu, Tai Wang, Wenwei Zhang, Runjian Chen, Jinkun Cao, Jiangmiao Pang, Dahua Lin
摘要
This paper introduces the Masked Voxel Jigsaw and Reconstruction (MV-JAR) method for LiDAR-based selfsupervised pre-training and a carefully designed dataefficient 3D object detection benchmark on the Waymo dataset. Inspired by the scene-voxel-point hierarchy in downstream 3D object detectors, we design masking and reconstruction strategies accounting for voxel distributions in the scene and local point distributions within the voxel. We employ a Reversed-Furthest-Voxel-Sampling strategy to address the uneven distribution of LiDAR points and propose MV-JAR, which combines two techniques for modeling the aforementioned distributions, resulting in superior performance. Our experiments reveal limitations in previous dataefficient experiments, which uniformly sample fine-tuning splits with varying data proportions from each LiDAR sequence, leading to similar data diversity across splits. To address this, we propose a new benchmark that samples scene sequences for diverse fine-tuning splits, ensuring adequate model convergence and providing a more accurate evaluation of pre-training methods. Experiments on our Waymo benchmark and the KITTI dataset demonstrate that MV-JAR consistently and significantly improves 3D detection performance across various data scales, achieving up to a 6.3% increase in mAPH compared to training from scratch. Codes and the benchmark will be available at https://github.com/SmartBot-PJLab/MV-JAR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- AD-PT: Autonomous Driving Pre-Training with Large-scale Point Cloud DatasetJiakang Yuan, Bo Zhang, Xiangchao Yan, Botian Shi 等NeurIPS 2023 · 被引用 36 次
- BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving ScenariosZhiwei Lin, Yongtao Wang, Shengxiang Qi, Nan Dong 等AAAI 2024 · 被引用 32 次
- UniPAD: A Universal Pre-Training Paradigm for Autonomous DrivingHonghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu 等CVPR 2024 · 被引用 31 次
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingChen Min, Dawei Zhao, Liang Xiao, Jian Zhao 等CVPR 2024 · 被引用 20 次
- MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object DetectionYuxue Yang, Lue Fan, Zhaoxiang ZhangICLR 2024 · 被引用 11 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Pre-training LiDAR-based 3D Object Detectors through ColorizationTai-Yu Pan, Chenyang Ma, Tianle Chen, Cheng Perng Phoo 等ICLR 2024 · 被引用 5 次
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 被引用 12 次
- Self-Supervised Representation Learning with Joint Embedding Predictive Architecture for Automotive LiDAR Object DetectionHaoran Zhu, Zhenyuan Dong, Kristi Topollai, Beiyao Sha 等AAAI 2026 · 被引用 5 次
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 被引用 333 次
- Exploring Geometry-aware Contrast and Clustering Harmonization for Self-supervised 3D Object DetectionHanxue Liang, Chenhan Jiang, Dapeng Feng, Xin Chen 等ICCV 2021 · 被引用 85 次
