Gradient Flossing: Improving Gradient Descent through Dynamic Control of Jacobians
Rainer Engelken
摘要
Training recurrent neural networks (RNNs) remains a challenge due to the instability of gradients across long time horizons, which can lead to exploding and vanishing gradients. Recent research has linked these problems to the values of Lyapunov exponents for the forward-dynamics, which describe the growth or shrinkage of infinitesimal perturbations. Here, we propose gradient flossing, a novel approach to tackling gradient instability by pushing Lyapunov exponents of the forward dynamics toward zero during learning. We achieve this by regularizing Lyapunov exponents through backpropagation using differentiable linear algebra. This enables us to "floss" the gradients, stabilizing them and thus improving network training. We demonstrate that gradient flossing controls not only the gradient norm but also the condition number of the long-term Jacobian, facilitating multidimensional error feedback propagation. We find that applying gradient flossing prior to training enhances both the success rate and convergence speed for tasks involving long time horizons. For challenging tasks, we show that gradient flossing during training can further increase the time horizon that can be bridged by backpropagation through time. Moreover, we demonstrate the effectiveness of our approach on various RNN architectures and tasks of variable temporal complexity. Additionally, we provide a simple implementation of our gradient flossing algorithm that can be used in practice. Our results indicate that gradient flossing via regularizing Lyapunov exponents can significantly enhance the effectiveness of RNN training and mitigate the exploding and vanishing gradients problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- A scalable generative model for dynamical system reconstruction from neuroimaging dataEric Volkmann, Alena Brändle, Daniel Durstewitz, Georgia KoppeNeurIPS 2024 · 被引用 14 次
- Predictability Enables Parallelization of Nonlinear State Space ModelsXavier Gonzalez, Leo Kozachkov, David M. Zoltowski, Kenneth L. Clarkson 等NeurIPS 2025 · 被引用 12 次
- Neural ODE Transformers: Analyzing Internal Dynamics and Adaptive Fine-tuningAnh Tong, Thanh Nguyen-Tang, Dongeun Lee, Duc Nguyen 等ICLR 2025
它引用的顶会 Paper3
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
- Coupled Oscillatory Recurrent Neural Network (coRNN): An accurate and (gradient) stable architecture for learning long time dependenciesT. Konstantin Rusch, Siddhartha MishraICLR 2021 · 被引用 121 次
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 被引用 109 次
相关 Paper
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- Time Adaptive Recurrent Neural NetworkAnil Kag, Venkatesh SaligramaCVPR 2021
- Unconditional stability of a recurrent neural circuit implementing divisive normalizationShivang Rawat, David J. Heeger, Stefano MartinianiNeurIPS 2024 · 被引用 8 次
- Training Recurrent Neural Networks via Forward Propagation Through TimeAnil Kag, Venkatesh SaligramaICML 2021 · 被引用 48 次
- Generalized Teacher Forcing for Learning Chaotic DynamicsFlorian Hess, Zahra Monfared, Manuel Brenner, Daniel DurstewitzICML 2023 · 被引用 67 次
