Training Recurrent Neural Networks via Forward Propagation Through Time
Anil Kag, Venkatesh Saligrama
摘要
Back-propagation through time (BPTT) has been widely used for training Recurrent Neural Networks (RNNs). BPTT updates RNN parameters on an instance by back-propagating the error in time over the entire sequence length, and as a result, leads to poor trainability due to the well-known gradient explosion/decay phenomena. While a number of prior works have proposed to mitigate vanishing/explosion effect through careful RNN architecture design, these RNN variants still train with BPTT. We propose a novel forwardpropagation algorithm, FPTT , where at each time, for an instance, we update RNN parameters by optimizing an instantaneous risk function. Our proposed risk is a regularization penalty at time t that evolves dynamically based on previously observed losses, and allows for RNN parameter updates to converge to a stationary solution of the empirical RNN objective. We consider both sequence-to-sequence as well as terminal loss problems. Empirically FPTT outperforms BPTT on a number of well-known benchmark tasks, thus enabling architectures like LSTMs to solve long range dependencies problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Online Training Through Time for Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He 等NeurIPS 2022 · 被引用 121 次
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 被引用 109 次
- Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural NetworksQingyan Meng, Mingqing Xiao, Shen Yan, Yisen Wang 等ICCV 2023 · 被引用 84 次
- NDOT: Neuronal Dynamics-based Online Training for Spiking Neural NetworksHaiyan Jiang, Giulia De Masi, Huan Xiong, Bin GuICML 2024 · 被引用 13 次
- Online Stabilization of Spiking Neural NetworksYaoyu Zhu, Jianhao Ding, Tiejun Huang, Xiaodong Xie 等ICLR 2024 · 被引用 13 次
它引用的顶会 Paper4
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 被引用 51 次
- Lipschitz Recurrent Neural NetworksN. Benjamin Erichson, Omri Azencot, Alejandro F. Queiruga, Liam Hodgkinson 等ICLR 2021 · 被引用 32 次
- Practical Real Time Recurrent Learning with a Sparse ApproximationJacob Menick, Erich Elsen, Utku Evci, Simon Osindero 等ICLR 2021 · 被引用 18 次
- Time Adaptive Recurrent Neural NetworkAnil Kag, Venkatesh SaligramaCVPR 2021
相关 Paper
- Training Recurrent Neural Networks Online by Learning Explicit State VariablesSomjit Nath, Vincent Liu, Alan Chan, Xin Li 等ICLR 2020 · 被引用 9 次
- Gradient Flossing: Improving Gradient Descent through Dynamic Control of JacobiansRainer EngelkenNeurIPS 2023 · 被引用 14 次
- Tuning the burn-in phase in training recurrent neural networks improves their performanceJulian D. Schiller, Malte Heinrich, Victor G. Lopez, Matthias A. MüllerICLR 2026 · 被引用 3 次
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
