SpikingLM: Towards Fully Spiking Language Model
Yu Liang, Zijian Zhou, Wenjie Wei, Shuai Wang, Honglin Cao, Ammar Belatreche, Yu Yang, Malu Zhang, Yang Yang, Haizhou Li
Abstract
Leveraging event-driven computation mechanism, Spiking Neural Networks (SNNs) have emerged as a representative paradigm for energy-efficient edge intelligence. However, extending SNNs to modern deep language models still faces two fundamental challenges. First, dead neurons in deep SNNs lead to degraded gradients, limiting the training effectiveness of spiking language models. Second, removing Softmax for energy efficiency weakens token-wise competition, reducing the model’s ability to select salient tokens. To address these challenges, we propose Spiking Language Model (SpikingLM) to bridge the efficiency of SNNs and the capability of modern language models through two key innovations. First, we propose Distribution-aware Scaling method, which rescales linear outputs into an activation-friendly range to alleviate dead neurons and stabilize gradient propagation. Notably, its scaling parameters can be fused into the preceding linear layers, incurring no additional inference overhead. Second, we introduce Spike2Max to restore winner-takes-all mechanism via base-2 exponentiation and max-subtraction. Compared with Softmax, Spike2Max reduces energy consumption by over 95% using hardware-efficient bit-shift operations. Experiments show that SpikingLM reduces energy consumption by 57.9% and achieves state-of-the-art performance on GLUE, laying a promising foundation for energy-efficient language modeling. Code is available at https://github.com/hamings1/SpikingLM.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6df78e4f-5e09-4476-9e30-cb8ca3e3177bBuilds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
- Going Deeper With Directly-Trained Larger Spiking Neural NetworksHanle Zheng, Yujie Wu, Lei Deng, Yifan Hu et al.AAAI 2021 · 694 citations
Related papers
- Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language ModelKaiwen Tang, Zhanglu Yan, Weng-Fai WongICML 2025
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao et al.ICML 2024 · 34 citations
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit DifferentiationMalyaban Bal, Abhronil SenguptaAAAI 2024 · 78 citations
- Energy-Efficient and Dequantization-Free Quantization of LLMs: A Spiking Neural Network Approach to Salient Value MitigationChenyu Wang, Zhanglu Yan, Zhi Zhou, Xu Chen et al.WWW 2026
- Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for TransformerYufei Guo, Xiaode Liu, Yuanpei Chen, Weihang Peng et al.CVPR 2025
