SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
Xingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao, Yiming Ju, Siqi Fan, Yequan Wang, Jiajun Zhang, Guoqi Li
Abstract
Towards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, making it a valuable issue to explore more general spike-driven models. However, the binary spikes in existing SNNs fail to encode adequate semantic information, placing technological challenges for generalization. This work proposes the first fully spiking mechanism for general language tasks, including both discriminative and generative ones. Different from previous spikes with 0,1 levels, we propose a more general spike formulation with bi-directional, elastic amplitude, and elastic frequency encoding, while still maintaining the addition nature of SNNs. In a single time step, the spike is enhanced by direction and amplitude information; in spike frequency, a strategy to control spike firing rate is well designed. We plug this elastic bi-spiking mechanism in language modeling, named SpikeLM. It is the first time to handle general language tasks with fully spike-driven models, which achieve much higher accuracy than previously possible. SpikeLM also greatly bridges the performance gap between SNNs and ANNs in language modeling. Our code is available at https://github.com/Xingrun- Xing/SpikeLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdfd2999-e39b-40d2-9690-dc14a0f160b1Cited by top-tier papers11
- Phi: Leveraging Pattern-based Hierarchical Sparsity for High-Efficiency Spiking Neural NetworksChiyue Wei, Bowen Duan, Cong Guo, Jingyang Zhang et al.ISCA 2025 · 9 citations
- Toward Relative Positional Encoding in Spiking TransformersChangze Lv, Yansen Wang, Dongqi Han, Yifei Shen et al.NeurIPS 2025 · 8 citations
- Positional Encoding for Spiking TransformersZijian Zhou, Yu Liang, Honglin Cao, Ammar Belatreche et al.ICML 2026 · 7 citations
- Spikingformer: A Key Foundation Model for Spiking Neural NetworksChenlin Zhou, Liutao Yu, Zhaokun Zhou, Han Zhang et al.AAAI 2026 · 4 citations
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command RecognitionJiaqi Wang, Liutao Yu, Xiongri Shen, Sihang Guo et al.AAAI 2026 · 1 citation
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Incorporating Learnable Membrane Time Constant to Enhance Learning of Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Timothée Masquelier et al.ICCV 2021 · 731 citations
Related papers
- SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based SpikingXingrun Xing, Boyan Gao, Zheng Liu, David A. Clifton et al.ICLR 2025
- Ternary Spike: Learning Ternary Spikes for Spiking Neural NetworksYufei Guo, Yuanpei Chen, Xiaode Liu, Weihang Peng et al.AAAI 2024 · 70 citations
- SpikingLM: Towards Fully Spiking Language ModelYu Liang, Zijian Zhou, Wenjie Wei, Shuai Wang et al.ICML 2026
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit DifferentiationMalyaban Bal, Abhronil SenguptaAAAI 2024 · 78 citations
- Spik-NeRF: Spiking Neural Networks for Neural Radiance FieldsGang Wan, Qinlong Lan, Zihan Li, Huimin Wang et al.NeurIPS 2025 · 1 citation
