Rethinking Spiking Self-Attention Mechanism: Implementing a-XNOR Similarity Calculation in Spiking Transformers
Yichen Xiao, Shuai Wang, Dehao Zhang, Wenjie Wei, Yimeng Shan, Xiaoli Liu, Yulin Jiang, Malu Zhang
摘要
Transformers significantly raise the performance limits across various tasks, spurring research into integrating them into spiking neural networks. However, a notable performance gap remains between existing spiking Transformers and their artificial neural network counterparts. Here, we first analyze the cause of this gap and attribute it to the dot product's ineffectiveness in measuring similarity between spiking queries and keys, due to numerous nonspiking events. To address this, we propose a novel α-XNOR similarity measure tailored for spike trains. It redefines the correlation between non-spike pairs as a specific value α, effectively overcoming the limitations of dot-product similarity. Furthermore, considering the sparse nature of spike trains where spikes carry more information than non-spikes, the α-XNOR similarity correspondingly highlights the distinct importance of spikes over non-spikes. Extensive experiments demonstrate that α-XNOR similarity significantly improves performance across different spiking Transformer architectures on various static and neuromorphic datasets, further revealing the potential of spiking Transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingDehao Zhang, Malu Zhang, Shuai Wang, Jingya Wang 等NeurIPS 2025 · 被引用 7 次
- Positional Encoding for Spiking TransformersZijian Zhou, Yu Liang, Honglin Cao, Ammar Belatreche 等ICML 2026 · 被引用 7 次
- Bipolar Self-attention for Spiking TransformersShuai Wang, Malu Zhang, Jingya Wang, Dehao Zhang 等NeurIPS 2025 · 被引用 4 次
- Training-Free ANN-to-SNN Conversion for High-Performance Spiking TransformersJingya Wang, Xin Deng, Wenjie Wei, Dehao Zhang 等AAAI 2026 · 被引用 1 次
- Robust Spiking Neural Networks Against Adversarial AttacksShuai Wang, Malu Zhang, Yulin Jiang, Dehao Zhang 等ICLR 2026
它引用的顶会 Paper29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
相关 Paper
- AdaS: Adaptive Gradient Descent for Spiking TransformersZijian Zhou, Honglin Cao, Ammar Belatreche, Wenjie Wei 等ICML 2026
- Rethinking Attention in Spiking Transformers: Overcoming Density Bias with Set SimilarityJinGyo Lim, Seunggyu Jeong, Seong-Eun KimICML 2026
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao 等CVPR 2025
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang 等ICLR 2023 · 被引用 103 次
- QKFormer: Hierarchical Spiking Transformer using Q-K AttentionChenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu 等NeurIPS 2024 · 被引用 126 次
