SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
Xingrun Xing, Boyan Gao, Zheng Liu, David A. Clifton, Shitao Xiao, Wanpeng Zhang, Li Du, Zheng Zhang, Guoqi Li, Jiajun Zhang
摘要
Recent advancements in large language models (LLMs) with billions of parameters have improved performance in various applications, but their inference processes demand significant energy and computational resources. In contrast, the human brain, with approximately 86 billion neurons, is much more energy-efficient than LLMs with similar parameters. Inspired by this, we redesign 7∼70 billion parameter LLMs using bio-plausible spiking mechanisms, emulating the efficient behavior of the human brain. We propose the first spiking large language model, SpikeLLM. Coupled with the proposed model, two essential approaches are proposed to improve spike training efficiency: Generalized Integrate-and-Fire (GIF) neurons to compress spike length from T to T L log 2 L bits, and an Optimal Brain Spiking framework to divide outlier channels and allocate different T for GIF neurons, which further compresses spike length to approximate log 2 T bits. The necessity of spike-driven LLM is proved by comparison with quantized LLMs with similar operations. In the OmniQuant pipeline, SpikeLLM reduces 11.01% WikiText2 perplexity and improves 2.55% accuracy of common scene reasoning on a LLAMA-7B W4A4 model. In the GPTQ pipeline, SpikeLLM achieves direct additive in linear layers, significantly exceeding PB-LLMs. Our code is publicly available at https://github.com/Xingrun-Xing2/SpikeLLM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Otters: An Energy-Efficient Spiking Transformer via Optical Time-to-First-Spike EncodingZhanglu Yan, Jiayi Mao, Qianhui Liu, Fanfan Li 等ICLR 2026 · 被引用 4 次
- Adaptive Fission: Post-training Encoding for Low-latency Spike Neural NetworksYizhou Jiang, Feng Chen, Yihan Li, Yuqian Liu 等NeurIPS 2025 · 被引用 2 次
- Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural NetworksJieyuan Zhang, Xiaolong Zhou, Shuai Wang, Wenjie Wei 等NeurIPS 2025 · 被引用 2 次
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command RecognitionJiaqi Wang, Liutao Yu, Xiongri Shen, Sihang Guo 等AAAI 2026 · 被引用 1 次
- SpikePack: Enhanced Information Flow in Spiking Neural Networks with High Hardware CompatibilityGuobin Shen, Jindong Li, Tenglong Li, Dongcheng Zhao 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
相关 Paper
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao 等ICML 2024 · 被引用 34 次
- Energy-Efficient and Dequantization-Free Quantization of LLMs: A Spiking Neural Network Approach to Salient Value MitigationChenyu Wang, Zhanglu Yan, Zhi Zhou, Xu Chen 等WWW 2026
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit DifferentiationMalyaban Bal, Abhronil SenguptaAAAI 2024 · 被引用 78 次
- Learn To be Efficient: Build Structured Sparsity in Large Language ModelsHaizhong Zheng, Xiaoyan Bai, Xueshen Liu, Zhuoqing Morley Mao 等NeurIPS 2024 · 被引用 29 次
- SpikingLM: Towards Fully Spiking Language ModelYu Liang, Zijian Zhou, Wenjie Wei, Shuai Wang 等ICML 2026
