LinK: Linear Kernel for LiDAR-based 3D Perception
Tao Lu, Xiang Ding, Haisong Liu, Gangshan Wu, Limin Wang
摘要
Extending the success of 2D Large Kernel to 3D perception is challenging due to: 1. the cubically-increasing overhead in processing 3D data; 2. the optimization difficulties from data scarcity and sparsity. Previous work has taken the first step to scale up the kernel size from 3 × 3 × 3 to 7 × 7 × 7 by introducing block-shared weights. However, to reduce the feature variations within a block, it only employs modest block size and fails to achieve larger kernels like the 21 × 21 × 21. To address this issue, we propose a new method, called LinK, to achieve a wider-range perception receptive field in a convolution-like manner with two core designs. The first is to replace the static kernel matrix with a linear kernel generator, which adaptively provides weights only for non-empty voxels. The second is to reuse the precomputed aggregation results in the overlapped blocks to reduce computation complexity. The proposed method successfully enables each voxel to perceive context within a range of 21 × 21 × 21. Extensive experiments on two basic perception tasks, 3D object detection and 3D semantic segmentation, demonstrate the effectiveness of our method. Notably, we rank 1st on the public leaderboard of the 3D detection benchmark of nuScenes (LiDAR track), by simply incorporating a LinK-based backbone into the basic detector, CenterPoint. We also boost the strong segmentation baseline's mIoU with 2.7% in the SemanticKITTI test set. Code is available at https://github.com/MCG-NJU/LinK .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SparseBEV: High-Performance Sparse 3D Object Detection from Multi-Camera VideosHaisong Liu, Yao Teng, Tao Lu, Haiguang Wang 等ICCV 2023 · 被引用 204 次
- Voxel Mamba: Group-Free State Space Models for Point Cloud based 3D Object DetectionGuowen Zhang, Lue Fan, Chenhang He, Zhen Lei 等NeurIPS 2024 · 被引用 137 次
- HEDNet: A Hierarchical Encoder-Decoder Network for 3D Object Detection in Point CloudsGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li 等NeurIPS 2023 · 被引用 95 次
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye 等NeurIPS 2024 · 被引用 84 次
- SAFDNet: A Simple and Effective Network for Fully Sparse 3D Object DetectionGang Zhang, Junnan Chen, Guohuan Gao, Jianmin Li 等CVPR 2024 · 被引用 56 次
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
相关 Paper
- LargeKernel3D: Scaling up Kernels in 3D Sparse CNNsYukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi 等CVPR 2023
- LSK3DNet: Towards Effective and Efficient 3D Perception with Large Sparse KernelsTuo Feng, Wenguan Wang, Fan Ma, Yi YangCVPR 2024
- Center-Based 3D Object Detection and TrackingTianwei Yin, Xingyi Zhou, Philipp KrähenbühlCVPR 2021
- Cylindrical and Asymmetrical 3D Convolution Networks for LiDAR SegmentationXinge Zhu, Hui Zhou, Tai Wang, Fangzhou Hong 等CVPR 2021
- Spherical Transformer for LiDAR-Based 3D RecognitionXin Lai, Yukang Chen, Fanbin Lu, Jianhui Liu 等CVPR 2023
