SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks
Xiubo Liang, Jinxing Han, Yuke Li, Haoqi Zhu, Yu Zhao, Hongzhi Wang
Abstract
Spiking Neural Networks (SNNs) enable eventdriven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2b0ad07-7e8b-4d40-a880-7914b64ce0aaBuilds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Deep Residual Learning in Spiking Neural NetworksWei Fang, Zhaofei Yu, Yanqi Chen, Tiejun Huang et al.NeurIPS 2021 · 857 citations
- Spike-driven TransformerMan Yao, Jiakui Hu, Zhaokun Zhou, Li Yuan et al.NeurIPS 2023 · 368 citations
- Enabling Deep Spiking Neural Networks with Hybrid Conversion and Spike Timing Dependent BackpropagationNitin Rathi, Gopalakrishnan Srinivasan, Priyadarshini Panda, Kaushik RoyICLR 2020 · 347 citations
Related papers
- Spiking Transformer with Experts MixtureZhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che et al.NeurIPS 2024 · 15 citations
- Spikformer: When Spiking Neural Network Meets TransformerZhaokun Zhou, Yuesheng Zhu, Chao He, Yaowei Wang et al.ICLR 2023 · 103 citations
- Spiking Transformer: Introducing Accurate Addition-Only Spiking Self-Attention for TransformerYufei Guo, Xiaode Liu, Yuanpei Chen, Weihang Peng et al.CVPR 2025
- Bipolar Self-attention for Spiking TransformersShuai Wang, Malu Zhang, Jingya Wang, Dehao Zhang et al.NeurIPS 2025 · 4 citations
- Towards High-performance Spiking Transformers from ANN to SNN ConversionZihan Huang, Xinyu Shi, Zecheng Hao, Tong Bu et al.ACM MM 2024 · 17 citations
