Exploiting Symmetric Temporally Sparse BPTT for Efficient RNN Training
Xi Chen, Chang Gao, Zuowen Wang, Longbiao Cheng, Sheng Zhou, Shih-Chii Liu, Tobi Delbruck
Abstract
Recurrent Neural Networks (RNNs) are useful in temporal sequence tasks. However, training RNNs involves dense matrix multiplications which require hardware that can support a large number of arithmetic operations and memory accesses. Implementing online training of RNNs on the edge calls for optimized algorithms for an efficient deployment on hardware. Inspired by the spiking neuron model, the Delta RNN exploits temporal sparsity during inference by skipping over the update of hidden states from those inactivated neurons whose change of activation across two timesteps is below a defined threshold. This work describes a training algorithm for Delta RNNs that exploits temporal sparsity in the backward propagation phase to reduce computational requirements for training on the edge. Due to the symmetric computation graphs of forward and backward propagation during training, the gradient computation of inactivated neurons can be skipped. Results show a reduction of ∼80% in matrix operations for training a 56k parameter Delta LSTM on the Fluent Speech Commands dataset with negligible accuracy loss. Logic simulations of a hardware accelerator designed for the training algorithm show 2-10X speedup in matrix computations for an activation sparsity range of 50%-90%. Additionally, we show that the proposed Delta RNN training will be useful for online incremental learning on edge devices with limited computing resources.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 080d5dd5-63d5-48cf-b359-0a226436f06cCited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured SparsityAlessandro Pierro, Steven Abreu, Jonathan Timcheck, Philipp Stratmann et al.ICML 2025
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingAnup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang et al.NeurIPS 2021 · 1 citation
- Selfish Sparse RNN TrainingShiwei Liu, Decebal Constantin Mocanu, Yulong Pei, Mykola PechenizkiyICML 2021 · 43 citations
- Spiking Neural Networks with Improved Inherent Recurrence Dynamics for Sequential LearningWachirawit Ponghiran, Kaushik RoyAAAI 2022 · 60 citations
- Practical Real Time Recurrent Learning with a Sparse ApproximationJacob Menick, Erich Elsen, Utku Evci, Simon Osindero et al.ICLR 2021 · 18 citations
