Synergy Between the Strong and the Weak: Spiking Neural Networks are Inherently Self-Distillers
Yongqi Ding, Lin Zuo, Mengmeng Jing, Kunshan Yang, Pei He, Tonglan Xie
Abstract
Brain-inspired spiking neural networks (SNNs) promise to be a low-power alternative to computationally intensive artificial neural networks (ANNs), although performance gaps persist. Recent studies have improved the performance of SNNs through knowledge distillation, but rely on large teacher models or introduce additional training overhead. In this paper, we show that SNNs can be naturally deconstructed into multiple submodels for efficient self-distillation. We treat each timestep instance of the SNN as a submodel and evaluate its output confidence, thus efficiently identifying the strong and the weak. Based on this strong and weak relationship, we propose two efficient self-distillation schemes: (1) Strong2Weak: During training, the stronger"teacher"guides the weaker"student", effectively improving overall performance. (2) Weak2Strong: The weak serve as the"teacher", distilling the strong in reverse with underlying dark knowledge, again yielding significant performance gains. For both distillation schemes, we offer flexible implementations such as ensemble, simultaneous, and cascade distillation. Experiments show that our method effectively improves the discriminability and overall performance of the SNN, while its adversarial robustness is also enhanced, benefiting from the stability brought by self-distillation. This ingeniously exploits the temporal properties of SNNs and provides insight into how to efficiently train high-performance SNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f258c816-8ac8-4d27-aca3-f0596d7acd2fCited by top-tier papers1
Ask how each one uses itBuilds on46
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
Related papers
- Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge DistillationQi Xu, Yaxin Li, Jiangrong Shen, Jian K. Liu et al.CVPR 2023
- Enhanced Self-Distillation Framework for Efficient Spiking Neural Network TrainingXiaochen Zhao, Chengting Yu, Kairong Yu, Lei Liu et al.NeurIPS 2025
- Efficient Logit-based Knowledge Distillation of Deep Spiking Neural Networks for Full-Range Timestep DeploymentChengting Yu, Xiaochen Zhao, Lei Liu, Shu Yang et al.ICML 2025
- Temporal Separation with Entropy Regularization for Knowledge Distillation in Spiking Neural NetworksKairong Yu, Chengting Yu, Tianqing Zhang, Xiaochen Zhao et al.CVPR 2025
- Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise ReplacementShu Yang, Chengting Yu, Lei Liu, Hanzhi Ma et al.CVPR 2025
