Spherical Transformer for LiDAR-Based 3D Recognition
Xin Lai, Yukang Chen, Fanbin Lu, Jianhui Liu, Jiaya Jia
摘要
LiDAR-based 3D point cloud recognition has benefited various applications. Without specially considering the Li-DAR point distribution, most current methods suffer from information disconnection and limited receptive field, especially for the sparse distant points. In this work, we study the varying-sparsity distribution of LiDAR points and present SphereFormer to directly aggregate information from dense close points to the sparse distant ones. We design radial window self-attention that partitions the space into multiple non-overlapping narrow and long windows. It overcomes the disconnection issue and enlarges the receptive field smoothly and dramatically, which significantly boosts the performance of sparse distant points. Moreover, to fit the narrow and long windows, we propose exponential splitting to yield fine-grained position encoding and dynamic feature selection to increase model representation ability. Notably, our method ranks 1 st on both nuScenes and SemanticKITTI semantic segmentation benchmarks with 81.9% and 74.8% mIoU, respectively. Also, we achieve the 3 rd place on nuScenes object detection benchmark with 72.8% NDS and 68.5% mAP. Code is available at https: // github.com/ dvlab-research/ SphereFormer.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- Perception-Aware Multi-Sensor Fusion for 3D LiDAR Semantic SegmentationZhuangwei Zhuang, Rong Li, Kui Jia, Qicheng Wang 等ICCV 2021 · 被引用 129 次
- Using a Waffle Iron for Automotive Point Cloud Semantic SegmentationGilles Puy, Alexandre Boulch, Renaud MarletICCV 2023 · 被引用 61 次
- Mask-Attention-Free Transformer for 3D Instance SegmentationXin Lai, Yuhui Yuan, Ruihang Chu, Yukang Chen 等ICCV 2023 · 被引用 53 次
- MCD: Diverse Large-Scale Multi-Campus Dataset for Robot PerceptionThien-Minh Nguyen, Shenghai Yuan, Thien Hoang Nguyen, Pengyu Yin 等CVPR 2024 · 被引用 48 次
- OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic SegmentationBohao Peng, Xiaoyang Wu, Li Jiang, Yukang Chen 等CVPR 2024 · 被引用 47 次
它引用的顶会 Paper43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- KPConv: Flexible and Deformable Convolution for Point CloudsHugues Thomas, Charles R. Qi, Jean-Emmanuel Deschaud, Beatriz Marcotegui 等ICCV 2019 · 被引用 3,193 次
相关 Paper
- OctFormer: Octree-based Transformers for 3D Point CloudsPeng-Shuai WangSIGGRAPH 2023 · 被引用 123 次
- Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR SegmentationXinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong 等CVPR 2021
- Efficient LiDAR Point Cloud Oversegmentation NetworkLe Hui, Linghua Tang, Yuchao Dai, Jin Xie 等ICCV 2023 · 被引用 7 次
- Rethinking Range View Representation for LiDAR SegmentationLingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma 等ICCV 2023 · 被引用 193 次
- Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic SegmentationYu Zheng, Guangming Wang, Jiuming Liu, Marc Pollefeys 等NeurIPS 2024 · 被引用 11 次
