Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
Zhenxin Lei, Man Yao, Jiakui Hu, Xinhao Luo, Yanye Lu, Bo Xu, Guoqi Li
Abstract
Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0× efficiency on ADE20K, +14.3% mIoU and 5.2× efficiency on VOC2012, and +9.1% mIoU and 6.6× efficiency on CityScapes. Our code is available at https://github.com/BICLab/Spike2Former
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7134f34-459a-439c-9240-252fbade37cdCited by top-tier papers10
- SpikeTrack: A Spike-driven Framework for Efficient Visual TrackingQiuyang Zhang, Jiujun Cheng, Qichao Mao, Cong Liu et al.CVPR 2026 · 7 citations
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural NetworksJieyuan Zhang, Xiaolong Zhou, Shuai Wang, Wenjie Wei et al.NeurIPS 2025 · 2 citations
- SpikeVLA: Vision-Language-Action Models with Spiking Neural NetworksRuiqi Song, Dujun Nie, Siyu Teng, Baiyong Ding et al.ICML 2026 · 1 citation
- TEFormer: Structured Bidirectional Temporal Enhancement Modeling in Spiking TransformersSicheng Shen, Mingyang Lv, Bing Han, Dongcheng Zhao et al.ICML 2026 · 1 citation
- ASG: Adaptive and Asymmetric Surrogate Gradients for Training Deep Spiking Neural NetworksYechan Kang, Yongjin Kweon, Mingyeong Seo, Sohee Park et al.ICML 2026
Builds on11
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Going Deeper With Directly-Trained Larger Spiking Neural NetworksHanle Zheng, Yujie Wu, Lei Deng, Yifan Hu et al.AAAI 2021 · 694 citations
- Optimal ANN-SNN Conversion for High-accuracy and Ultra-low-latency Spiking Neural NetworksTong Bu, Wei Fang, Jianhao Ding, Penglin Dai et al.ICLR 2022 · 272 citations
Related papers
- Adaptive Calibration: A Unified Conversion Framework of Spiking Neural NetworksZiqing Wang, Yuetong Fang, Jiahang Cao, Hongwei Ren et al.AAAI 2025 · 9 citations
- SpikeConverter: An Efficient Conversion Framework Zipping the Gap between Artificial Neural Networks and Spiking Neural NetworksFangxin Liu, Wenbo Zhao, Yongbiao Chen, Zongwu Wang et al.AAAI 2022 · 50 citations
- Inference-Scale Complexity in ANN-SNN Conversion for High-Performance and Low-Power ApplicationsTong Bu, Maohua Li, Zhaofei YuCVPR 2025
- Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic ChipsMan Yao, Jiakui Hu, Tianxiang Hu, Yifan Xu et al.ICLR 2024 · 154 citations
- SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and O(T) ComplexityShihao Zou, Qingfeng Li, Wei Ji, Jingjing Li et al.ICML 2025
