SmoothSpike: Spiking Transformer with Learnable Hadamard Transformation
Zijian Zhou, Wenjie Wei, Yu Liang, Jialin Li, Ammar Belatreche, Honglin Cao, Shuai Wang, Malu Zhang, Yang Yang, Haizhou Li
摘要
Spiking Neural Networks (SNNs) that leverage sparse binary spikes and temporal dynamics have emerged as energy-efficient alternatives to Artificial Neural Networks (ANNs). However, SNNs suffer from limited representational capacity due to the discrete nature of spikes. Existing solutions extending spike levels often overlook the constraints of the simulation time window, leading to a critical issue we identify as spike saturation-induced information homogenization. In this phenomenon, distinct high-amplitude inputs result in identical maximized spike counts, truncating the dynamic range and hindering the model’s ability to capture fine-grained semantic differences. To address this, we propose SmoothSpike, a novel method designed to enhance representational capacity by suppressing spike saturation. We first introduce a randomized Hadamard transformation to smooth neuronal inputs, theoretically proving its efficacy in constraining extreme values and reducing both saturation probability and input variability among saturated neurons. To further improve adaptability, we evolve this into a learnable orthogonal transformation. Initialized with Hadamard matrices and maintained orthogonal via Newton-Schulz iteration, this module dynamically adapts to varying input distributions during training. Extensive experiments on language modeling tasks show that SmoothSpike effectively mitigates the information homogenization problem and improves task performance. This positions SmoothSpike as a robust solution to bridge the performance gap between SNNs and ANNs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice CodebooksAlbert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov 等ICML 2024 · 被引用 295 次
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan 等ICLR 2022 · 被引用 121 次
- Parallel Spiking Neurons with High Efficiency and Ability to Learn Long-term DependenciesWei Fang, Zhaofei Yu, Zhaokun Zhou, Ding Chen 等NeurIPS 2023 · 被引用 104 次
- BiT: Robustly Binarized Multi-distilled TransformerZechun Liu, Barlas Oguz, Aasish Pappu, Lin Xiao 等NeurIPS 2022 · 被引用 93 次
相关 Paper
- SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsXingrun Xing, Zheng Zhang, Ziyi Ni, Shitao Xiao 等ICML 2024 · 被引用 34 次
- EnOF-SNN: Training Accurate Spiking Neural Networks via Enhancing the Output FeatureYufei Guo, Weihang Peng, Xiaode Liu, Yuanpei Chen 等NeurIPS 2024 · 被引用 21 次
- Training High-Performance Low-Latency Spiking Neural Networks by Differentiation on Spike RepresentationQingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang 等CVPR 2022 · 被引用 114 次
- Enhancing Representation of Spiking Neural Networks via Similarity-Sensitive Contrastive LearningYuhan Zhang, Xiaode Liu, Yuanpei Chen, Weihang Peng 等AAAI 2024 · 被引用 20 次
- DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural NetworksWei Liu, Li Yang, Mingxuan Zhao, Shuxun Wang 等ICLR 2025
