SVT-Net: Super Light-Weight Sparse Voxel Transformer for Large Scale Place Recognition
Zhaoxin Fan, Zhenbo Song, Hongyan Liu, Zhiwu Lu, Jun He, Xiaoyong Du
摘要
Simultaneous Localization and Mapping (SLAM) and Autonomous Driving are becoming increasingly more important in recent years. Point cloud-based large scale place recognition is the spine of them. While many models have been proposed and have achieved acceptable performance by learning short-range local features, they always skip long-range contextual properties. Moreover, the model size also becomes a serious shackle for their wide applications. To overcome these challenges, we propose a super light-weight network model termed SVT-Net. On top of the highly efficient 3D Sparse Convolution (SP-Conv), an Atom-based Sparse Voxel Transformer (ASVT) and a Cluster-based Sparse Voxel Transformer (CSVT) are proposed respectively to learn both short-range local features and long-range contextual features. Consisting of ASVT and CSVT, SVT-Net can achieve state-of-the-art performance in terms of both recognition accuracy and running speed with a super-light model size (0.9M parameters). Meanwhile, for the purpose of further boosting efficiency, we introduce two simplified versions, which also achieve state-of-the-art performance and further reduce the model size to 0.8M and 0.4M respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- CASSPR: Cross Attention Single Scan Place RecognitionYan Xia, Mariia Gladkova, Rui Wang, Qianyun Li 等ICCV 2023 · 被引用 72 次
- CrossLoc3D: Aerial-Ground Cross-Source 3D Place RecognitionTianrui Guan, Aswath Muthuselvam, Montana Hoover, Xijun Wang 等ICCV 2023 · 被引用 24 次
- TransLoc4D: Transformer-Based 4D Radar Place RecognitionGuohao Peng, Heshan Li, Yangyang Zhao, Jun Zhang 等CVPR 2024 · 被引用 19 次
- VLM-Loc: Localization in Point Cloud Maps via Vision-Language ModelsShuhao Kang, Youqi Liao, Peijie Wang, Wenlong Liao 等CVPR 2026 · 被引用 4 次
- L2RSI: Cross-view LiDAR-based Place Recognition for Large-scale Urban Scenes via Remote Sensing ImageryZiwei Shi, Xiaoran Zhang, Wenjing Xu, Yan Xia 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment AnalysisZhe Liu, Shunbo Zhou, Chuanzhe Suo, Peng Yin 等ICCV 2019 · 被引用 337 次
相关 Paper
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai 等ICCV 2021 · 被引用 535 次
- PVT-SSD: Single-Stage 3D Object Detector with Point-Voxel TransformerHonghui Yang, Wenxiao Wang, Minghao Chen, Binbin Lin 等CVPR 2023
- Pyramid Point Cloud Transformer for Large-Scale Place RecognitionLe Hui, Hang Yang, Mingmei Cheng, Jin Xie 等ICCV 2021 · 被引用 147 次
- Fast Point TransformerChunghyun Park, Yoonwoo Jeong, Minsu Cho, Jaesik ParkCVPR 2022
- DSVT: Dynamic Sparse Voxel Transformer with Rotated SetsHaiyang Wang, Chen Shi, Shaoshuai Shi, Meng Lei 等CVPR 2023
