Energy-Efficient and Dequantization-Free Quantization of LLMs: A Spiking Neural Network Approach to Salient Value Mitigation
Chenyu Wang, Zhanglu Yan, Zhi Zhou, Xu Chen, Weng-Fai Wong
Abstract
In the era of large language models (LLMs), weight-activation quantization helps fit models on edge device by reducing memory and compute bit-widths. However, three challenges persist for energy constrained hardware: (1) even after quantization, multiply-accumulate (MAC) operations remain unavoidable and continue to dominate energy consumption; (2) dequantization (or per-tensor/ channel rescaling) introduces extra arithmetic and data movement, increasing latency and energy; (3) uniform parameters bit widths clip salient values—while intra-channel mixed precision is generally impractical on current matrix hardware and memory. In contrast, brain-inspired Spiking Neural Networks (SNNs), owing to their binary spike-based information representation and the Integrate-and-Fire (IF) paradigm, naturally support mixed-precision storage and energy-efficient computation by replacing complex MACs with temporal Accumulate (ACCs). Motivated by this property, we propose SpikeQuant, which selectively applies mixed-precision quantization to activations with salient values and re-encodes them into binary spike counts, thereby enabling dynamic mixed storage of different bitwidths. Furthermore, by embedding the quantization scale into the threshold of the IF mechanism, our approach performs energy-efficient linear transformations on weights and activations while avoiding explicit dequantization. Experimental results demonstrate that SpikeQuant consistently achieves near-FP16 perplexity under W4A4 quantization while reducing energy cost by up to 4.6× compared to existing methods, highlighting its effectiveness for accurate and energy-efficient LLM deployment. Our code is open-sourced at https://github.com/wangchenyu929/SpikeQuant.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 59b23633-ed11-4a96-a56c-d07faccce44aBuilds on15
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
- Quantizable Transformers: Removing Outliers by Helping Attention Heads Do NothingYelysei Bondarenko, Markus Nagel, Tijmen BlankevoortNeurIPS 2023 · 196 citations
Related papers
- QP-SNN: Quantized and Pruned Spiking Neural NetworksWenjie Wei, Malu Zhang, Zijian Zhou, Ammar Belatreche et al.ICLR 2025
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao et al.ACM MM 2024 · 23 citations
- HardF-SNN: Hardware-Friendly Quantization for Spiking Neural Networks with Efficient Integer-Arithmetic-Only InferenceHanwen Liu, Kexin Shi, Jieyuan Zhang, Yimeng Shan et al.AAAI 2026
- Are Conventional SNNs Really Efficient? A Perspective from Network QuantizationGuobin Shen, Dongcheng Zhao, Tenglong Li, Jindong Li et al.CVPR 2024 · 8 citations
- Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inferenceKe Yi, Zengke Liu, Jianwei Zhang, Chengyuan Li et al.ICLR 2025
