3DSMT: A Hybrid Spiking Mamba-Transformer for Point Cloud Analysis
Zhiming Zhou, Yong He, Qiaoyun Wu, Chaoxu Mu, Ajmal Mian
Abstract
The sparse unordered structure of point clouds causes unnecessary computation and energy consumption in deep models. Conventionally, the Transformer architecture is leveraged to model global relationships in point clouds, however, its quadratic complexity restricts scalability. Although the Mamba architecture enables efficient global modeling with linear complexity, it lacks natural adaptability to unordered point clouds. Spiking Neural Network (SNN) is an energy-efficient alternative to Artificial Neural Network (ANN), offering an ultra low-power event-driven paradigm. The inherent sparsity and event-driven characteristics of SNN are highly compatible with the sparse distribution of point clouds. To balance efficiency and performance, we propose a hybrid spiking Mamba-Transformer (3DSMT) model for point cloud analysis. 3DSMT integrates a Spiking Local Offset Attention module to efficiently capture fine-grained local geometric features with a spiking Mamba block designed for unordered point clouds to achieve global feature integration with linear complexity. Experiments show that 3DSMT achieves state-of-the-art performance among SNN-based methods in shape classification, few-shot classification, and part segmentation tasks, significantly reducing computational energy consumption while also outperforming numerous ANN-based models. Our source code is in supplementary material and will be made publicly available
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a71de22-ab86-4646-badb-9f43eac53522Builds on19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World DataMikaela Angelina Uy, Quang-Hieu Pham, Binh-Son Hua, Duc Thanh Nguyen et al.ICCV 2019 · 1,003 citations
- Point Transformer V2: Grouped Vector Attention and Partition-based PoolingXiaoyang Wu, Yixing Lao, Li Jiang, Xihui Liu et al.NeurIPS 2022 · 924 citations
Related papers
- SpikeNet: Sparse Spike-Driven Mask Vector Transformer for Energy-Efficient and Stable Spiking Point Cloud ProcessingZhouzhiming Zhou, Yong He, Chaoxu Mu, Qiaoyun Wu et al.ICML 2026
- Spiking Point Transformer for Point Cloud ClassificationPeixi Wu, Bosong Chai, Hebei Li, Menghua Zheng et al.AAAI 2025 · 13 citations
- Spiking PointNet: Spiking Neural Networks for Point CloudsDayong Ren, Zhe Ma, Yuanpei Chen, Weihang Peng et al.NeurIPS 2023 · 64 citations
- Spiking Discrepancy Transformer for Point Cloud AnalysisYijie Lu, Zhiyi Pan, Renrui Zhang, Yanhao Jia et al.ICLR 2026
- Efficient 3D Recognition with Event-driven Spike Sparse ConvolutionXuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou et al.AAAI 2025 · 17 citations
