Practical Real Time Recurrent Learning with a Sparse Approximation
Jacob Menick, Erich Elsen, Utku Evci, Simon Osindero, Karen Simonyan, Alex Graves
摘要
Recurrent neural networks are usually trained with backpropagation through time, which requires storing a complete history of network states, and prohibits updating the weights "online" (after every timestep). Real Time Recurrent Learning (RTRL) eliminates the need for history storage and allows for online weight updates, but does so at the expense of computational costs that are quartic in the state size. This renders RTRL training intractable for all but the smallest networks, even ones that are made highly sparse. We introduce the Sparse n-step Approximation (SnAp) to the RTRL influence matrix. SnAp only tracks the influence of a parameter on hidden units that are reached by the computation graph within timesteps of the recurrent core. SnAp with is no more expensive than backpropagation but allows training on arbitrarily long sequences. We find that it substantially outperforms other RTRL approximations with comparable costs such as Unbiased Online Recurrent Optimization. For highly sparse networks, SnAp with remains tractable and can outperform backpropagation through time in terms of learning speed when updates are done online.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Online Training Through Time for Spiking Neural NetworksMingqing Xiao, Qingyan Meng, Zongpeng Zhang, Di He 等NeurIPS 2022 · 被引用 121 次
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 被引用 109 次
- Training Recurrent Neural Networks via Forward Propagation Through TimeAnil Kag, Venkatesh SaligramaICML 2021 · 被引用 48 次
- Exploring the Promise and Limits of Real-Time Recurrent LearningKazuki Irie, Anand Gopalakrishnan, Jürgen SchmidhuberICLR 2024 · 被引用 23 次
- Real-Time Recurrent Learning using Trace Units in Reinforcement LearningEsraa Elelimy, Adam White, Michael Bowling, Martha WhiteNeurIPS 2024 · 被引用 15 次
它引用的顶会 Paper2
相关 Paper
- Training Recurrent Neural Networks Online by Learning Explicit State VariablesSomjit Nath, Vincent Liu, Alan Chan, Xin Li 等ICLR 2020 · 被引用 9 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- RNNs Incrementally Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?Anil Kag, Ziming Zhang, Venkatesh SaligramaICLR 2020 · 被引用 51 次
- Explain My Surprise: Learning Efficient Long-Term Memory by predicting uncertain outcomesArtyom Y. Sorokin, Nazar Buzun, Leonid Pugachev, Mikhail BurtsevNeurIPS 2022 · 被引用 13 次
- Skipper: Enabling efficient SNN training through activation-checkpointing and time-skippingSonali Singh, Anup Sarma, Sen Lu, Abhronil Sengupta 等MICRO 2022 · 被引用 13 次
