Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for Transformer
Yufei Guo, Xiaode Liu, Yuanpei Chen, Weihang Peng, Yuhan Zhang, Zhe Ma
Abstract
Transformers have demonstrated outstanding performance across a wide range of tasks, owing to their selfattention mechanism, but they are highly energy-consuming. Spiking Neural Networks have emerged as a promising energy-efficient alternative to traditional Artificial Neural Networks, leveraging event-driven computation and binary spikes for information transfer. The combination of Transformers' capabilities with the energy efficiency of SNNs offers a compelling opportunity. This paper addresses the challenge of adapting the self-attention mechanism of Transformers to the spiking paradigm by introducing a novel approach: Accurate Addition-Only Spiking Self-Attention (A 2 OS 2 A). Unlike existing methods that rely solely on binary spiking neurons for all components of the self-attention mechanism, our approach integrates binary, ReLU, and ternary spiking neurons. This hybrid strategy significantly improves accuracy while preserving non-multiplicative computations. Moreover, our method eliminates the need for softmax and scaling operations. Extensive experiments show that the A 2 OS 2 A-based Spiking Transformer outperforms existing SNN-based Transformers on several datasets, even achieving an accuracy of 78.66% on ImageNet-1K. Our work represents a significant advancement in SNN-based Transformer models, offering a more accurate and efficient solution for real-world applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a09041d-bca8-4ef1-bf37-ab73e9267f1aCited by top-tier papers7
- MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksDengfeng Xue, Wenjuan Li, Yifan Lu, Chunfeng Yuan et al.NeurIPS 2025
- Robust Spiking Neural Networks by Temporal Mutual InformationMengting Xu, Shi Gu, Peng Lin, De Ma et al.CVPR 2026
- ReverB-SNN: Reversing Bit of the Weight and Activation for Spiking Neural NetworksYufei Guo, Yuhan Zhang, Jie Zhou, Xiaode Liu et al.ICML 2025
- Temporal Interaction in Spiking Transformers with Multi-Delay MixerKexin Shi, Hanwen Liu, Zeyang Song, Yang Liu et al.CVPR 2026
- Kronecker Generative Networks: A General Neural Architecture for Parameter-Efficient Learning Across Classification TasksYang Yang, Zhengmin Kong, Yuan Liu, Tao Huang et al.ICML 2026
Builds on33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- Bipolar Self-attention for Spiking TransformersShuai Wang, Malu Zhang, Jingya Wang, Dehao Zhang et al.NeurIPS 2025 · 4 citations
- SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural NetworksXinyu Shi, Zecheng Hao, Zhaofei YuCVPR 2024 · 53 citations
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao et al.CVPR 2025
