SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition
Zhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu, Jun He, Xiaoyong Du
Abstract
Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range local features, they always skip long-range contextual properties. Moreover, the model size also becomes a serious shackle for their wide applications. To overcome these challenges, we propose a super light-weight network model termed SVT-Net. On top of the highly efficient 3D Sparse Convolution (SP-Conv), an Atom-based Sparse Voxel Transformer (ASVT) and a Cluster-based Sparse Voxel Transformer (CSVT) are proposed respectively to learn both short-range local features and long-range contextual features. Consisting of ASVT and CSVT, SVT-Net can achieve state-of-the-art performance in terms of both recognition accuracy and running speed with a super-light model size (0.9M parameters). Meanwhile, for the purpose of further boosting efficiency, we introduce two simplified versions, which also achieve state-of-the-art performance and further reduce the model size to 0.8M and 0.4M respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4fb54d5-156a-415f-86fb-b372f71b7b34Cited by top-tier papers10
- CASSPR: Cross Attention Single Scan Place RecognitionYan Xia, Mariia Gladkova, Rui Wang, Qianyun Li et al.ICCV 2023 · 72 citations
- CrossLoc3D: Aerial-Ground Cross-Source 3D Place RecognitionTianrui Guan, Aswath Muthuselvam, Montana Hoover, Xijun Wang et al.ICCV 2023 · 24 citations
- TransLoc4D: Transformer-Based 4D Radar Place RecognitionGuohao Peng, Heshan Li, Yangyang Zhao, Jun Zhang et al.CVPR 2024 · 19 citations
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao et al.CVPR 2026 · 4 citations
- L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing ImageryZiwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia et al.NeurIPS 2025 · 3 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment AnalysisZhe Liu, Shunbo Zhou, Chuanzhe Suo, Peng Yin et al.ICCV 2019 · 337 citations
Related papers
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel TransformerHonghui Yang, Wenxiao Wang, Minghao Chen, Binbin Lin et al.CVPR 2023
- Pyramid Point Cloud Transformer for Large-Scale Place RecognitionLe Hui, Hang Yang, Mingmei Cheng, Jin Xie et al.ICCV 2021 · 147 citations
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
- DSVT: Dynamic Sparse Voxel Transformer with Rotated SetsHaiyang Wang, Chen Shi, Shaoshuai Shi, Meng Lei et al.CVPR 2023
