SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation
Zijian Zhou, Wenjie Wei, Yu Liang, Jialin Li, Ammar Belatreche, Honglin Cao, Shuai Wang, Malu Zhang, Yang Yang, Haizhou Li
Abstract
Spiking Neural Networks (SNNs) that leverage sparse binary spikes and temporal dynamics have emerged as energy-efficient alternatives to Artificial Neural Networks (ANNs). However, SNNs suffer from limited representational capacity due to the discrete nature of spikes. Existing solutions extending spike levels often overlook the constraints of the simulation time window, leading to a critical issue we identify as spike saturation-induced information homogenization. In this phenomenon, distinct high-amplitude inputs result in identical maximized spike counts, truncating the dynamic range and hindering the model’s ability to capture fine-grained semantic differences. To address this, we propose SmoothSpike, a novel method designed to enhance representational capacity by suppressing spike saturation. We first introduce a randomized Hadamard transformation to smooth neuronal inputs, theoretically proving its efficacy in constraining extreme values and reducing both saturation probability and input variability among saturated neurons. To further improve adaptability, we evolve this into a learnable orthogonal transformation. Initialized with Hadamard matrices and maintained orthogonal via Newton-Schulz iteration, this module dynamically adapts to varying input distributions during training. Extensive experiments on language modeling tasks show that SmoothSpike effectively mitigates the information homogenization problem and improves task performance. This positions SmoothSpike as a robust solution to bridge the performance gap between SNNs and ANNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ccca6826-7bdc-4ca4-8324-5764872fc52dBuilds on18
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov et al.ICML 2024 · 295 citations
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan et al.ICLR 2022 · 121 citations
- Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term DependenciesWei Fang, Zhaofei Yu, Zhaokun Zhou, Ding Chen et al.NeurIPS 2023 · 104 citations
- BiT: Robustly Binarized Multi-distilled TransformerZechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao et al.NeurIPS 2022 · 93 citations
Related papers
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao et al.ICML 2024 · 34 citations
- EnOF-SNN: Training Accurate Spiking Neural Networks via Enhancing the Output FeatureYufei Guo, Weihang Peng, Xiaode Liu, Yuanpei Chen et al.NeurIPS 2024 · 21 citations
- Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike RepresentationQingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang et al.CVPR 2022 · 114 citations
- Enhancing Representation of Spiking Neural Networks via Similarity-Sensitive Contrastive LearningYuhan Zhang, Xiaode Liu, Yuanpei Chen, Weihang Peng et al.AAAI 2024 · 20 citations
- DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural NetworksWei Liu, Li Yang, Mingxuan Zhao, Shuxun Wang et al.ICLR 2025
