RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving
Angelika Ando, Spyros Gidaris, Andrei Bursuc, Gilles Puy, Alexandre Boulch, Renaud Marlet
Abstract
Casting semantic segmentation of outdoor LiDAR point clouds as a 2D problem, e.g., via range projection, is an effective and popular approach. These projection-based methods usually benefit from fast computations and, when combined with techniques which use other point cloud representations, achieve state-of-the-art results. Today, projection-based methods leverage 2D CNNs but recent advances in computer vision show that vision transformers (ViTs) have achieved state-of-the-art results in many imagebased benchmarks. In this work, we question if projectionbased methods for 3D semantic segmentation can benefit from these latest improvements on ViTs. We answer positively but only after combining them with three key ingredients: (a) ViTs are notoriously hard to train and require a lot of training data to learn powerful representations. By preserving the same backbone architecture as for RGB images, we can exploit the knowledge from long training on large image collections that are much cheaper to acquire and annotate than point clouds. We reach our best results with pre-trained ViTs on large image datasets. (b) We compensate ViTs' lack of inductive bias by substituting a tailored convolutional stem for the classical linear embedding layer. (c) We refine pixel-wise predictions with a convolutional decoder and a skip connection from the convolutional stem to combine low-level but fine-grained features of the the convolutional stem with the high-level but coarse predictions of the ViT encoder. With these ingredients, we show that our method, called RangeViT, outperforms existing projection-based methods on nuScenes and SemanticKITTI. The code is available at https:// github.com/valeoai/rangevit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- Rethinking Range View Representation for LiDAR SegmentationLingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma et al.ICCV 2023 · 193 citations
- UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg CodebaseYouquan Liu, Runnan Chen, Xin Li, Lingdong Kong et al.ICCV 2023 · 94 citations
- Using a Waffle Iron for Automotive Point Cloud Semantic SegmentationGilles Puy, Alexandre Boulch, Renaud MarletICCV 2023 · 61 citations
- Is Your LiDAR Placement Optimized for 3D Scene Understanding?Ye Li, Lingdong Kong, Hanjiang Hu, Xiaohao Xu et al.NeurIPS 2024 · 36 citations
- Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic SegmentationYu Zheng, Guangming Wang, Jiuming Liu, Marc Pollefeys et al.NeurIPS 2024 · 11 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- Masked Clustering Prediction for Unsupervised Point Cloud Pre-trainingBin Ren, Xiaoshui Huang, Mengyuan Liu, Hong Liu et al.AAAI 2026 · 1 citation
- Self-Supervised Pretraining for Large-Scale Point CloudsZaiwei Zhang, Min Bai, Li Erran LiNeurIPS 2022 · 12 citations
- A Simple Vision Transformer for Weakly Semi-supervised 3D Object DetectionDingyuan Zhang, Dingkang Liang, Zhikang Zou, Jingyu Li et al.ICCV 2023 · 36 citations
- RegFormer: An Efficient Projection-Aware Transformer Network for Large-Scale Point Cloud RegistrationJiuming Liu, Guangming Wang, Zhe Liu, Chaokang Jiang et al.ICCV 2023 · 71 citations
- Efficient 3D Semantic Segmentation with Superpoint TransformerDamien Robert, Hugo Raguet, Loïc LandrieuICCV 2023 · 131 citations
