Image-to-Lidar Self-Supervised Distillation for Autonomous Driving Data
Corentin Sautier, Gilles Puy, Spyros Gidaris, Alexandre Boulch, Andrei Bursuc, Renaud Marlet
摘要
Segmenting or detecting objects in sparse Lidar point clouds are two important tasks in autonomous driving to allow a vehicle to act safely in its 3D environment. The best performing methods in 3D semantic segmentation or object detection rely on a large amount of annotated data. Yet annotating 3D Lidar data for these tasks is tedious and costly. In this context, we propose a self-supervised pretraining method for 3D perception models that is tailored to autonomous driving data. Specifically, we leverage the availability of synchronized and calibrated image and Lidar sensors in autonomous driving setups for distilling self-supervised pre-trained image representations into 3D models. Hence, our method does not require any point cloud nor image annotations. The keyingredient of our method is the use of superpixels which are used to pool 3D point features and 2D pixel features in visually similar regions. We then train a 3D network on the self-supervised task of matching these pooled point features with the corresponding pooled image pixel features. The advantages of contrasting regions obtained by superpixels are that: (1) grouping together pixels and points of visually coherent regions leads to a more meaningful contrastive task that produces features well adapted to 3D semantic segmentation and 3D object detection; (2) all the different regions have the same weight in the contrastive loss regardless of the number of 3D points sampled in these regions; (3) it mitigates the noise produced by incorrect matching of points and pixels due to occlusions between the different sensors. Extensive experiments on autonomous driving datasets demonstrate the ability of our image-to-Lidar distillation strategy to produce 3D representations that transfer well on semantic segmentation and object detection tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper53
- Segment Any Point Cloud Sequences by Distilling Vision Foundation ModelsYouquan Liu, Lingdong Kong, Jun Cen, Runnan Chen 等NeurIPS 2023 · 被引用 169 次
- Towards Label-free Scene Understanding by Vision Foundation ModelsRunnan Chen, Youquan Liu, Lingdong Kong, Nenglun Chen 等NeurIPS 2023 · 被引用 82 次
- Bridging the Domain Gap: Self-Supervised 3D Scene Understanding with Foundation ModelsZhimin Chen, Longlong Jing, Yingwei Li, Bing LiNeurIPS 2023 · 被引用 57 次
- Unsupervised 3D Perception with 2D Vision-Language Distillation for Autonomous DrivingMahyar Najibi, Jingwei Ji, Yin Zhou, Charles R. Qi 等ICCV 2023 · 被引用 52 次
- VLM2Scene: Self-Supervised Image-Text-LiDAR Learning with Foundation Models for Autonomous Driving Scene UnderstandingGuibiao Liao, Jiankun Li, Xiaoqing YeAAAI 2024 · 被引用 46 次
它引用的顶会 Paper32
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
相关 Paper
- ALSO: Automotive Lidar Self-Supervision by Occupancy EstimationAlexandre Boulch, Corentin Sautier, Björn Michele, Gilles Puy 等CVPR 2023
- Implicit Surface Contrastive Clustering for LiDAR Point CloudsZaiwei Zhang, Min Bai, Li Erran LiCVPR 2023
- Temporal Consistent 3D LiDAR Representation Learning for Semantic Perception in Autonomous DrivingLucas Nunes, Louis Wiesmann, Rodrigo Marcuzzi, Xieyuanli Chen 等CVPR 2023
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 被引用 12 次
- Three Pillars Improving Vision Foundation Model Distillation for LidarGilles Puy, Spyros Gidaris, Alexandre Boulch, Oriane Siméoni 等CVPR 2024 · 被引用 19 次
