LinK: Linear Kernel for LiDAR-based 3D Perception
Tao Lu, Xiang Ding, Haisong Liu, Gangshan Wu, Limin Wang
Abstract
Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3 × 3 × 3 to 7 × 7 × 7 by introducing block-shared weights. However, to reduce the feature variations within a block, it only employs modest block size and fails to achieve larger kernels like the 21 × 21 × 21. To address this issue, we propose a new method, called LinK, to achieve a wider-range perception receptive field in a convolution-like manner with two core designs. The first is to replace the static kernel matrix with a linear kernel generator, which adaptively provides weights only for non-empty voxels. The second is to reuse the precomputed aggregation results in the overlapped blocks to reduce computation complexity. The proposed method successfully enables each voxel to perceive context within a range of 21 × 21 × 21. Extensive experiments on two basic perception tasks, 3D object detection and 3D semantic segmentation, demonstrate the effectiveness of our method. Notably, we rank 1st on the public leaderboard of the 3D detection benchmark of nuScenes (LiDAR track), by simply incorporating a LinK-based backbone into the basic detector, CenterPoint. We also boost the strong segmentation baseline's mIoU with 2.7% in the SemanticKITTI test set. Code is available at https://github.com/MCG-NJU/LinK .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b53d4bc-6810-467f-a776-e6859ddd8d8eCited by top-tier papers12
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang et al.ICCV 2023 · 204 citations
- Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object DetectionGuowen Zhang, Lue Fan, Chenhang He, Zhen Lei et al.NeurIPS 2024 · 137 citations
- HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point CloudsGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li et al.NeurIPS 2023 · 95 citations
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye et al.NeurIPS 2024 · 84 citations
- SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object DetectionGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li et al.CVPR 2024 · 56 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
Related papers
- LargeKernel3D: Scaling up Kernels in 3D Sparse CNNsYukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi et al.CVPR 2023
- LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse KernelsTuo Feng, Wenguan Wang, Fan Ma, Yi YangCVPR 2024
- Center-Based 3D Object Detection and TrackingTianwei Yin, Xingyi Zhou, Philipp KrähenbühlCVPR 2021
- Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR SegmentationXinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong et al.CVPR 2021
- Spherical Transformer for LiDAR-Based 3D RecognitionXin Lai, Yukang Chen, Fanbin Lu, Jianhui Liu et al.CVPR 2023
