Identifying Equivalent Training Dynamics
William T. Redman, Juan M. Bello-Rivas, Maria Fonoberova, Ryan Mohr, Yannis G. Kevrekidis, Igor Mezic
Abstract
Study of the nonlinear evolution deep neural network (DNN) parameters undergo during training has uncovered regimes of distinct dynamical behavior. While a detailed understanding of these phenomena has the potential to advance improvements in training efficiency and robustness, the lack of methods for identifying when DNN models have equivalent dynamics limits the insight that can be gained from prior work. Topological conjugacy, a notion from dynamical systems theory, provides a precise definition of dynamical equivalence, offering a possible route to address this need. However, topological conjugacies have historically been challenging to compute. By leveraging advances in Koopman operator theory, we develop a framework for identifying conjugate and non-conjugate training dynamics. To validate our approach, we demonstrate that comparing Koopman eigenvalues can correctly identify a known equivalence between online mirror descent and online gradient descent. We then utilize our approach to: (a) identify non-conjugate training dynamics between shallow and wide fully connected neural networks; (b) characterize the early phase of training dynamics in convolutional neural networks; (c) uncover non-conjugate training dynamics in Transformers that do and do not undergo grokking. Our results, across a range of DNN architectures, illustrate the flexibility of our framework and highlight its potential for shedding new light on training dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e03cf9c6-c6d3-4940-903d-828d0decdedfCited by top-tier papers8
- InputDSA: Demixing, then comparing recurrent and externally driven dynamicsAnn Huang, Mitchell Ostrow, Satpreet H. Singh, Leo Kozachkov et al.ICLR 2026 · 9 citations
- Not so griddy: Internal representations of RNNs path integrating more than one agentWilliam Redman, Francisco Acosta, Santiago Acosta-Mendoza, Nina MiolaneNeurIPS 2024 · 8 citations
- A Geometry-Aware Metric for Mode Collapse in Time Series Generative ModelsYassine Abbahaddou, Amine Mohamed AboussalahNeurIPS 2025 · 2 citations
- A Spectral-Grassmann Wasserstein metric for operator representations of dynamical systemsThibaut Germain, Rémi Flamary, Vladimir R Kostic, Karim LouniciICLR 2026 · 2 citations
- Temporal Geometry of Deep Networks: Hyperbolic Representations of Training Dynamics for Intrinsic ExplainabilityAmbarish MoharilICLR 2026
Builds on15
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
- Optimizing Neural Networks via Koopman Operator TheoryAkshunna S. Dogra, William T. RedmanNeurIPS 2020 · 65 citations
- Beyond Geometry: Comparing the Temporal Structure of Computation in Neural Circuits with Dynamical Similarity AnalysisMitchell Ostrow, Adam Eisen, Leo Kozachkov, Ila FieteNeurIPS 2023 · 60 citations
Related papers
- SKOLR: Structured Koopman Operator Linear RNN for Time-Series ForecastingYitian Zhang, Liheng Ma, Antonios Valkanas, Boris N. Oreshkin et al.ICML 2025
- An Operator Theoretic Approach for Analyzing Sequence Neural NetworksIlan Naiman, Omri AzencotAAAI 2023 · 14 citations
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 86 citations
- Progress measures for grokking via mechanistic interpretabilityNeel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith et al.ICLR 2023 · 54 citations
- Decoupling Dynamical Richness from Representation Learning: Towards Practical MeasurementYoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard et al.ICLR 2026 · 2 citations
