SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit Differentiation
Malyaban Bal, Abhronil Sengupta
Abstract
Large language Models (LLMs), though growing exceedingly powerful, comprises of orders of magnitude less neurons and synapses than the human brain. However, it requires significantly more power/energy to operate. In this work, we propose a novel bio-inspired spiking language model (LM) which aims to reduce the computational cost of conventional LMs by drawing motivation from the synaptic information flow in the brain. In this paper, we demonstrate a framework that leverages the average spiking rate of neurons at equilibrium to train a neuromorphic spiking LM using implicit differentiation technique, thereby overcoming the non-differentiability problem of spiking neural network (SNN) based algorithms without using any type of surrogate gradient. The steady-state convergence of the spiking neurons also allows us to design a spiking attention mechanism, which is critical in developing a scalable spiking LM. Moreover, the convergence of average spiking rate of neurons at equilibrium is utilized to develop a novel ANN-SNN knowledge distillation based technique wherein we use a pre-trained BERT model as “teacher” to train our “student” spiking architecture. While the primary architecture proposed in this paper is motivated by BERT, the technique can be potentially extended to different kinds of LLMs. Our work is the first one to demonstrate the performance of an operational spiking LM architecture on multiple different tasks in the GLUE benchmark. Our implementation source code is available at https://github.com/NeuroCompLab-psu/SpikingBERT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52d0207c-553f-4130-b7f1-c4375ea5d8c6Cited by top-tier papers21
- Autonomous Driving with Spiking Neural NetworksRuijie Zhu, Ziqing Wang, Leilani Gilpin, Jason EshraghianNeurIPS 2024 · 35 citations
- Rethinking the Membrane Dynamics and Optimization Objectives of Spiking Neural NetworksHangchi Shen, Qian Zheng, Huamin Wang, Gang PanNeurIPS 2024 · 30 citations
- SpikedAttention: Training-Free and Fully Spike-Driven Transformer-to-SNN Conversion with Winner-Oriented Spike Shift for Softmax OperationSangwoo Hwang, Seunghyun Lee, Dahoon Park, Donghun Lee et al.NeurIPS 2024 · 23 citations
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao et al.ACM MM 2024 · 23 citations
- Spiking Transformer with Experts MixtureZhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che et al.NeurIPS 2024 · 15 citations
Builds on6
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- Training Feedback Spiking Neural Networks by Implicit Differentiation on the Equilibrium StateMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Yisen Wang et al.NeurIPS 2021 · 83 citations
- M-FAC: Efficient Matrix-Free Approximations of Second-Order InformationElias Frantar, Eldar Kurtic, Dan AlistarhNeurIPS 2021 · 69 citations
- NAS-BERT: Task-Agnostic and Adaptive-Size BERT Compression with Neural Architecture SearchJin Xu, Xu Tan, Renqian Luo, Kaitao Song et al.KDD 2021 · 49 citations
- The Optimal BERT Surgeon: Scalable and Accurate Second-Order Pruning for Large Language ModelsEldar Kurtic, Daniel Campos, Tuan Nguyen, Elias Frantar et al.EMNLP 2022 · 4 citations
Related papers
- Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge DistillationQi Xu, Yaxin Li, Jiangrong Shen, Jian K. Liu et al.CVPR 2023
- Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language ModelKaiwen Tang, Zhanglu Yan, Weng-Fai WongICML 2025
- SpikingLM: Towards Fully Spiking Language ModelYu Liang, Zijian Zhou, Wenjie Wei, Shuai Wang et al.ICML 2026
- Efficient ANN-Guided Distillation: Aligning Rate-based Features of Spiking Neural Networks through Hybrid Block-wise ReplacementShu Yang, Chengting Yu, Lei Liu, Hanzhi Ma et al.CVPR 2025
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao et al.ICML 2024 · 34 citations
