Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers
Yichen Xiao, Shuai Wang, Dehao Zhang, Wenjie Wei, Yimeng Shan, Xiaoli Liu, Yulin Jiang, Malu Zhang
Abstract
Transformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyze the cause of this gap and attribute it to the dot product's ineffectiveness in measuring similarity between spiking queries and keys, due to numerous nonspiking events. To address this, we propose a novel α-XNOR similarity measure tailored for spike trains. It redefines the correlation between non-spike pairs as a specific value α, effectively overcoming the limitations of dot-product similarity. Furthermore, considering the sparse nature of spike trains where spikes carry more information than non-spikes, the α-XNOR similarity correspondingly highlights the distinct importance of spikes over non-spikes. Extensive experiments demonstrate that α-XNOR similarity significantly improves performance across different spiking Transformer architectures on various static and neuromorphic datasets, further revealing the potential of spiking Transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 475a4645-76fc-4e6e-af04-72bb86e22b56Cited by top-tier papers10
- Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingDehao Zhang, Malu Zhang, Shuai Wang, Jingya Wang et al.NeurIPS 2025 · 7 citations
- Positional Encoding for Spiking TransformersZijian Zhou, Yu Liang, Honglin Cao, Ammar Belatreche et al.ICML 2026 · 7 citations
- Bipolar Self-attention for Spiking TransformersShuai Wang, Malu Zhang, Jingya Wang, Dehao Zhang et al.NeurIPS 2025 · 4 citations
- Training-Free ANN-to-SNN Conversion for High-Performance Spiking TransformersJingya Wang, Xin Deng, Wenjie Wei, Dehao Zhang et al.AAAI 2026 · 1 citation
- Robust Spiking Neural Networks Against Adversarial AttacksShuai Wang, Malu Zhang, Yulin Jiang, Dehao Zhang et al.ICLR 2026
Builds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
Related papers
- AdaS: Adaptive Gradient Descent for Spiking TransformersZijian Zhou, Honglin Cao, Ammar Belatreche, Wenjie Wei et al.ICML 2026
- Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set SimilarityJinGyo Lim, Seunggyu Jeong, Seong-Eun KimICML 2026
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao et al.CVPR 2025
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- QKFormer: Hierarchical Spiking Transformer using Q-K AttentionChenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu et al.NeurIPS 2024 · 126 citations
