Bifurcations and loss jumps in RNN training
Lukas Eisenmann, Zahra Monfared, Niclas Alexander Göring, Daniel Durstewitz
摘要
Recurrent neural networks (RNNs) are popular machine learning tools for modeling and forecasting sequential data and for inferring dynamical systems (DS) from observed time series. Concepts from DS theory (DST) have variously been used to further our understanding of both, how trained RNNs solve complex tasks, and the training process itself. Bifurcations are particularly important phenomena in DS, including RNNs, that refer to topological (qualitative) changes in a system's dynamical behavior as one or more of its parameters are varied. Knowing the bifurcation structure of an RNN will thus allow to deduce many of its computational and dynamical properties, like its sensitivity to parameter variations or its behavior during training. In particular, bifurcations may account for sudden loss jumps observed in RNN training that could severely impede the training process. Here we first mathematically prove for a particular class of ReLU-based RNNs that certain bifurcations are indeed associated with loss gradients tending toward infinity or zero. We then introduce a novel heuristic algorithm for detecting all fixed points and k-cycles in ReLU-based RNNs and their existence and stability regions, hence bifurcation manifolds in parameter space. In contrast to previous numerical algorithms for finding fixed points and common continuation methods, our algorithm provides exact results and returns fixed points and cycles up to high orders with surprisingly good scaling behavior. We exemplify the algorithm on the analysis of the training process of RNNs, and find that the recently introduced technique of generalized teacher forcing completely avoids certain types of bifurcations in training. Thus, besides facilitating the DST analysis of trained RNNs, our algorithm provides a powerful instrument for analyzing the training process itself.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Inferring stochastic low-rank recurrent neural networks from neural dataMatthijs Pals, A Erdem Sagtekin, Felix Pei, Manuel Glöckler 等NeurIPS 2024 · 被引用 37 次
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term StatisticsChristoph Jürgen Hemmer, Daniel DurstewitzNeurIPS 2025 · 被引用 25 次
- Almost-Linear RNNs Yield Highly Interpretable Symbolic Codes in Dynamical Systems ReconstructionManuel Brenner, Christoph Jürgen Hemmer, Zahra Monfared, Daniel DurstewitzNeurIPS 2024 · 被引用 22 次
- Integrating Multimodal Data for Joint Generative Modeling of Complex DynamicsManuel Brenner, Florian Hess, Georgia Koppe, Daniel DurstewitzICML 2024 · 被引用 18 次
- Optimal Recurrent Network Topologies for Dynamical Systems ReconstructionChristoph Jürgen Hemmer, Manuel Brenner, Florian Hess, Daniel DurstewitzICML 2024 · 被引用 6 次
它引用的顶会 Paper12
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- On the difficulty of learning chaotic dynamics with RNNsJonas M. Mikhaeil, Zahra Monfared, Daniel DurstewitzNeurIPS 2022 · 被引用 109 次
- Generalized Teacher Forcing for Learning Chaotic DynamicsFlorian Hess, Zahra Monfared, Manuel Brenner, Daniel DurstewitzICML 2023 · 被引用 67 次
- Tractable Dendritic RNNs for Reconstructing Nonlinear Dynamical SystemsManuel Brenner, Florian Hess, Jonas M. Mikhaeil, Leonard F. Bereska 等ICML 2022 · 被引用 48 次
- Reverse engineering recurrent neural networks with Jacobian switching linear dynamical systemsJimmy T. H. Smith, Scott W. Linderman, David SussilloNeurIPS 2021 · 被引用 44 次
相关 Paper
- Detecting Invariant Manifolds in ReLU-Based RNNsLukas Eisenmann, Alena Brändle, Zahra Monfared, Daniel DurstewitzICLR 2026 · 被引用 3 次
- Topology and geometry of the learning space of ReLU networks: connectivity and singularitiesMarco Nurisso, Pierrick Leroy, Giovanni Petri, Francesco VaccarinoICLR 2026 · 被引用 6 次
- Identifying nonlinear dynamical systems with multiple time scales and long-range dependenciesDominik Schmidt, Georgia Koppe, Zahra Monfared, Max Beutelspacher 等ICLR 2021 · 被引用 41 次
- Continuous-Time Piecewise-Linear Recurrent Neural NetworksAlena Brändle, Lukas Eisenmann, Florian Götz, Daniel DurstewitzICML 2026 · 被引用 2 次
- On Logical Extrapolation for Mazes with Recurrent and Implicit NetworksBrandon Knutson, Amandin Chyba Rabeendran, Michael I. Ivanitskiy, Jordan Pettyjohn 等AAAI 2026 · 被引用 7 次
