Training Recurrent Neural Networks via Forward Propagation Through Time
Anil Kag, Venkatesh Saligrama
Abstract
Back-propagation through time (BPTT) has been widely used for training Recurrent Neural Networks (RNNs). BPTT updates RNN parameters on an instance by back-propagating the error in time over the entire sequence length, and as a result, leads to poor trainability due to the well-known gradient explosion/decay phenomena. While a number of prior works have proposed to mitigate vanishing/explosion effect through careful RNN architecture design, these RNN variants still train with BPTT. We propose a novel forwardpropagation algorithm, FPTT , where at each time, for an instance, we update RNN parameters by optimizing an instantaneous risk function. Our proposed risk is a regularization penalty at time t that evolves dynamically based on previously observed losses, and allows for RNN parameter updates to converge to a stationary solution of the empirical RNN objective. We consider both sequence-to-sequence as well as terminal loss problems. Empirically FPTT outperforms BPTT on a number of well-known benchmark tasks, thus enabling architectures like LSTMs to solve long range dependencies problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5342442e-b88d-43ea-bd53-bf94ff83af75Cited by top-tier papers10
- Online Training Through Time for Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He et al.NeurIPS 2022 · 121 citations
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 109 citations
- Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural NetworksQingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang et al.ICCV 2023 · 84 citations
- NDOT: Neuronal Dynamics-based Online Training for Spiking Neural NetworksHaiyan Jiang, Giulia De Masi, Huan Xiong, Bin GuICML 2024 · 13 citations
- Online Stabilization of Spiking Neural NetworksYaoyu Zhu, Jianhao Ding, Tiejun Huang, Xiaodong Xie et al.ICLR 2024 · 13 citations
Builds on4
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 51 citations
- Lipschitz Recurrent Neural NetworksN. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson et al.ICLR 2021 · 32 citations
- Practical Real Time Recurrent Learning with a Sparse ApproximationJacob Menick, Erich Elsen, Utku Evci, Simon Osindero et al.ICLR 2021 · 18 citations
- Time Adaptive Recurrent Neural NetworkAnil Kag, Venkatesh SaligramaCVPR 2021
Related papers
- Training Recurrent Neural Networks Online by Learning Explicit State VariablesSomjit Nath, Vincent Liu, Alan Chan, Xin Li et al.ICLR 2020 · 9 citations
- Gradient Flossing: Improving Gradient Descent through Dynamic Control of JacobiansRainer EngelkenNeurIPS 2023 · 14 citations
- Tuning the burn-in phase in training recurrent neural networks improves their performanceJulian D. Schiller, Malte Heinrich, Victor G. Lopez, Matthias A. MüllerICLR 2026 · 3 citations
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher et al.ICLR 2021 · 41 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
