MsSVT: Mixed-scale Sparse Voxel Transformer for 3D Object Detection on Point Clouds
Shaocong Dong, Lihe Ding, Haiyang Wang, Tingfa Xu, Xinli Xu, Jie Wang, Ziyang Bian, Ying Wang, Jianan Li
Abstract
3D object detection from the LiDAR point cloud is fundamental to autonomous driving. Large-scale outdoor scenes usually feature significant variance in instance scales, thus requiring features rich in long-range and fine-grained information to support accurate detection. Recent detectors leverage the power of window-based transformers to model long-range dependencies but tend to blur out fine-grained details. To mitigate this gap, we present a novel Mixed-scale Sparse Voxel Trans-former, named MsSVT, which can well capture both types of information simultaneously by the divide-and-conquer philosophy. Specifically, MsSVT explicitly divides attention heads into multiple groups, each in charge of attending to information within a particular range. All groups’ output is merged to obtain the final mixed-scale features. Moreover, we provide a novel chessboard sampling strategy to reduce the computational complexity of applying a window-based transformer in 3D voxel space. To improve efficiency, we also implement the voxel sampling and gathering operations sparsely with a hash map. Endowed by the powerful capability and high efficiency of modeling mixed-scale information, our single-stage detector built on top of MsSVT surprisingly outperforms state-of-the-art two-stage detectors on Waymo. Our project page: https://github.com/dscdyc/MsSVT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b0915dfb-7b8c-4bc3-8458-512392a0f7deCited by top-tier papers10
- UniTR: A Unified and Efficient Multi-Modal Transformer for Bird's-Eye-View RepresentationHaiyang Wang, Hao Tang, Shaoshuai Shi, Aoxue Li et al.ICCV 2023 · 106 citations
- LION: Linear Group RNN for 3D Object Detection in Point CloudsZhe Liu, Jinghua Hou, Xinyu Wang, Xiaoqing Ye et al.NeurIPS 2024 · 84 citations
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao et al.NeurIPS 2023 · 65 citations
- Clusterformer: Cluster-based Transformer for 3D Object Detection in Point CloudsYu Pei, Xian Zhao, Hao Li, Jingyuan Ma et al.ICCV 2023 · 13 citations
- SwiftPillars: High-Efficiency Pillar Encoder for Lidar-Based 3D DetectionXin Jin, Kai Liu, Cong Ma, Ruining Yang et al.AAAI 2024 · 12 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- How Do Vision Transformers Work?Namuk Park, Songkuk KimICLR 2022 · 653 citations
Related papers
- Embracing Single Stride 3D Object Detector with Sparse TransformerLue Fan, Ziqi Pang, Tianyuan Zhang, Yu-Xiong Wang et al.CVPR 2022
- PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel TransformerHonghui Yang, Wenxiao Wang, Minghao Chen, Binbin Lin et al.CVPR 2023
- Fully Sparse 3D Object DetectionLue Fan, Feng Wang, Naiyan Wang, Zhaoxiang ZhangNeurIPS 2022 · 168 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- GeoFormer: Geometry Point Encoder for 3D Object Detection with Graph-Based TransformerXin Jin, Haisheng Su, Cong Ma, Kai Liu et al.ICCV 2025 · 2 citations
