Towards Efficient Spiking Transformer: a Token Sparsification Framework for Training and Inference Acceleration
Zhengyang Zhuge, Peisong Wang, Xingting Yao, Jian Cheng
Abstract
Nowadays Spiking Transformers have exhibited remarkable performance close to Artificial Neural Networks (ANNs), while enjoying the inherent energy-efficiency of Spiking Neural Networks (SNNs). However, training Spiking Transformers on GPUs is considerably more timeconsuming compared to the ANN counterparts, despite the energy-efficient inference through neuromorphic computation. In this paper, we investigate the token sparsification technique for efficient training of Spiking Transformer and find conventional methods suffer from noticeable performance degradation. We analyze the issue and propose our Sparsification with Timestep-wise Anchor Token and dual Alignments (STATA). Timestep-wise Anchor Token enables precise identification of important tokens across timesteps based on standardized criteria. Additionally, dual Alignments incorporate both Intra and Inter Alignment of the attention maps, fostering the learning of inferior attention. Extensive experiments show the effectiveness of STATA thoroughly, which demonstrates up to ∼1.53× training speedup and ∼48% energy reduction with comparable performance on various datasets and architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- TP-Spikformer: Token Pruned Spiking TransformerWenjie Wei, Xiaolong Zhou, Malu Zhang, Ammar Belatreche et al.ICLR 2026 · 6 citations
- MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksDengfeng Xue, Wenjuan Li, Yifan Lu, Chunfeng Yuan et al.NeurIPS 2025
- Efficient Parallel Training Methods for Spiking Neural Networks with Constant Time ComplexityWanjin Feng, Xingyu Gao, Wenqian Du, Hailong Shi et al.ICML 2025
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Going Deeper With Directly-Trained Larger Spiking Neural NetworksHanle Zheng, Yujie Wu, Lei Deng, Yifan Hu et al.AAAI 2021 · 694 citations
Related papers
- Spiking Transformer with Spatial-Temporal AttentionDonghyun Lee, Yuhang Li, Youngeun Kim, Shiting Xiao et al.CVPR 2025
- TPipe: Efficient Spiking Transformer Training with Time Parallelism and Asynchronous PipelineYubing Bao, ZhiHui Lu, Qiang Duan, Changze Lv et al.INFOCOM 2026 · 1 citation
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- SpikingResformer: Bridging ResNet and Vision Transformer in Spiking Neural NetworksXinyu Shi, Zecheng Hao, Zhaofei YuCVPR 2024 · 53 citations
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
