Harnessing Manycore Processors with Distributed Memory for Accelerated Training of Sparse and Recurrent Models
Jan Finkbeiner, Thomas Gmeinder, Mark Pupilli, Alexander Titterton, Emre Neftci
Abstract
Current AI training infrastructure is dominated by single instruction multiple data (SIMD) and systolic array architectures, such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs), that excel at accelerating parallel workloads and dense vector matrix multiplications. Potentially more efficient neural network models utilizing sparsity and recurrence cannot leverage the full power of SIMD processor and are thus at a severe disadvantage compared to today's prominent parallel architectures like Transformers and CNNs, thereby hindering the path towards more sustainable AI. To overcome this limitation, we explore sparse and recurrent model training on a massively parallel multiple instruction multiple data (MIMD) architecture with distributed local memory. We implement a training routine based on backpropagation though time (BPTT) for the brain-inspired class of Spiking Neural Networks (SNNs) that feature binary sparse activations. We observe a massive advantage in using sparse activation tensors with a MIMD processor, the Intelligence Processing Unit (IPU) compared to GPUs. On training workloads, our results demonstrate 5-10x throughput gains compared to A100 GPUs and up to 38x gains for higher levels of activation sparsity, without a significant slowdown in training convergence or reduction in final model performance. Furthermore, our results show highly promising trends for both single and multi IPU configurations as we scale up to larger model sizes. Our work paves the way towards more efficient, non-standard models via AI training hardware beyond GPUs, and competitive large scale SNN models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f26c72f4-9d78-4658-8082-3efba76636d4Builds on4
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Sparse Spiking Gradient DescentNicolas Perez Nieves, Dan F. M. GoodmanNeurIPS 2021 · 105 citations
- Shfl-BW: accelerating deep neural network inference with tensor-core aware weight pruningGuyue Huang, Haoran Li, Minghai Qin, Fei Sun et al.DAC 2022 · 19 citations
- Efficient recurrent architectures through activity sparsity and sparse back-propagation through timeAnand Subramoney, Khaleelulla Khan Nazeer, Mark Schöne, Christian Mayr et al.ICLR 2023 · 7 citations
Related papers
- Skipper: Enabling efficient SNN training through activation-checkpointing and time-skippingSonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta et al.MICRO 2022 · 13 citations
- SpinalFlow: An Architecture and Dataflow Tailored for Spiking Neural NetworksSurya Narayanan, Karl Taht, Rajeev Balasubramonian, Edouard Giacomin et al.ISCA 2020 · 122 citations
- Multiplication-Free Parallelizable Spiking Neurons with Efficient Spatio-Temporal DynamicsPeng Xue, Wei Fang, Zhengyu Ma, Zihan Huang et al.NeurIPS 2025 · 5 citations
- High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference CostJiakui Hu, Man Yao, Xuerui Qiu, Yuhong Chou et al.ICML 2024 · 24 citations
- Prosperity: Accelerating Spiking Neural Networks via Product SparsityChiyue Wei, Cong Guo, Feng Cheng, Shiyu Li et al.HPCA 2025 · 14 citations
