Spherical Transformer for LiDAR-Based 3D Recognition
Xin Lai, Yukang Chen, Fanbin Lu, Jianhui Liu, Jiaya Jia
Abstract
LiDAR-based 3D point cloud recognition has benefited various applications. Without specially considering the Li-DAR point distribution, most current methods suffer from information disconnection and limited receptive field, especially for the sparse distant points. In this work, we study the varying-sparsity distribution of LiDAR points and present SphereFormer to directly aggregate information from dense close points to the sparse distant ones. We design radial window self-attention that partitions the space into multiple non-overlapping narrow and long windows. It overcomes the disconnection issue and enlarges the receptive field smoothly and dramatically, which significantly boosts the performance of sparse distant points. Moreover, to fit the narrow and long windows, we propose exponential splitting to yield fine-grained position encoding and dynamic feature selection to increase model representation ability. Notably, our method ranks 1 st on both nuScenes and SemanticKITTI semantic segmentation benchmarks with 81.9% and 74.8% mIoU, respectively. Also, we achieve the 3 rd place on nuScenes object detection benchmark with 72.8% NDS and 68.5% mAP. Code is available at https: // github.com/ dvlab-research/ SphereFormer.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers66
- Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic SegmentationZhuangwei Zhuang, Rong Li, Kui Jia, Qicheng Wang et al.ICCV 2021 · 129 citations
- Using a Waffle Iron for Automotive Point Cloud Semantic SegmentationGilles Puy, Alexandre Boulch, Renaud MarletICCV 2023 · 61 citations
- Mask-Attention-Free Transformer for 3D Instance SegmentationXin Lai, Yuhui Yuan, Ruihang Chu, Yukang Chen et al.ICCV 2023 · 53 citations
- MCD: Diverse Large-Scale Multi-Campus Dataset for Robot PerceptionThien-Minh Nguyen, Shenghai Yuan, Thien Hoang Nguyen, Pengyu Yin et al.CVPR 2024 · 48 citations
- OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic SegmentationBohao Peng, Xiaoyang Wu, Li Jiang, Yukang Chen et al.CVPR 2024 · 47 citations
Builds on43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui et al.ICCV 2019 · 3,193 citations
Related papers
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 123 citations
- Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR SegmentationXinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong et al.CVPR 2021
- Efficient LiDAR Point Cloud Oversegmentation NetworkLe Hui, Linghua Tang, Yuchao Dai, Jin Xie et al.ICCV 2023 · 7 citations
- Rethinking Range View Representation for LiDAR SegmentationLingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma et al.ICCV 2023 · 193 citations
- Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic SegmentationYu Zheng, Guangming Wang, Jiuming Liu, Marc Pollefeys et al.NeurIPS 2024 · 11 citations
