LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
Zeyu Liu, Gourav Datta, Anni Li, Peter Anthony Beerel
Abstract
Transformer models have demonstrated high accuracy in numerous applications but have high complexity and lack sequential processing capability making them ill-suited for many streaming applications at the edge where devices are heavily resource-constrained. Thus motivated, many researchers have proposed reformulating the transformer models as RNN modules which modify the self-attention computation with explicit states. However, these approaches often incur significant performance degradation. The ultimate goal is to develop a model that has the following properties: parallel training, streaming and low-cost inference, and state-of-the-art (SOTA) performance. In this paper, we propose a new direction to achieve this goal. We show how architectural modifications to a fully-sequential recurrent model can help push its performance toward Transformer models while retaining its sequential processing capability. Specifically, inspired by the recent success of Legendre Memory Units (LMU) in sequence learning tasks, we propose LMUFormer, which augments the LMU with convolutional patch embedding and convolutional channel mixer. Moreover, we present a spiking version of this architecture, which introduces the benefit of states within the patch embedding and channel mixer modules while simultaneously reducing the computing complexity. We evaluated our architectures on multiple sequence datasets. Of particular note is our performance on the Speech Commands V2 dataset (35 classes). In comparison to SOTA transformer-based models within the ANN domain, our LMUFormer demonstrates comparable performance while necessitating a remarkable 53× reduction in parameters and a substantial 65× decrement in FLOPs. Furthermore, when benchmarked against extant low-complexity SNN variants, our model establishes a new SOTA with an accuracy of 96.12%. Additionally, owing to our model's proficiency in real-time data processing, we are able to achieve a 32.03% reduction in sequence length, all while incurring an inconsequential decline in performance. Our code is publicly available at https://github.com/zeyuliu1037/LMUFormer.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab0d233e-c812-48c0-9fde-2a7df338ddb3Cited by top-tier papers7
- SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space ModelsShuaijie Shen, Chao Wang, Renzhuo Huang, Yan Zhong et al.AAAI 2025 · 22 citations
- Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingDehao Zhang, Malu Zhang, Shuai Wang, Jingya Wang et al.NeurIPS 2025 · 7 citations
- P-Spikessm: Harnessing Probabilistic Spiking State Space Models for Long-Range Dependency TasksMalyaban Bal, Abhronil SenguptaICLR 2025 · 3 citations
- SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command RecognitionJiaqi Wang, Liutao Yu, Xiongri Shen, Sihang Guo et al.AAAI 2026 · 1 citation
- LIF Recurrent Memory Enables Long-Horizon Spiking ComputationFenghao Liu, Yipeng Shen, Peng Chen, Qian Zheng et al.ICML 2026
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
Related papers
- Parallelizing Legendre Memory Unit TrainingNarsimha Reddy Chilkuri, Chris EliasmithICML 2021 · 47 citations
- QKFormer: Hierarchical Spiking Transformer using Q-K AttentionChenlin Zhou, Han Zhang, Zhaokun Zhou, Liutao Yu et al.NeurIPS 2024 · 126 citations
- SpiLiFormer: Enhancing Spiking Transformers with Lateral InhibitionZeqi Zheng, Yanchen Huang, Yingchao Yu, Zizheng Zhu et al.ICCV 2025 · 2 citations
- Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured SparsityAlessandro Pierro, Steven Abreu, Jonathan Timcheck, Philipp Stratmann et al.ICML 2025
- EdgeFormer: A Parameter-Efficient Transformer for On-Device Seq2seq GenerationTao Ge, Si-Qing Chen, Furu WeiEMNLP 2022 · 16 citations
