The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial Conditions
Gül Sena Altintas, Devin Kwok, Colin Raffel, David Rolnick
Abstract
Neural network training is inherently sensitive to initialization and the randomness induced by stochastic gradient descent. However, it is unclear to what extent such effects lead to meaningfully different networks, either in terms of the models’ weights or the underlying functions that were learned. In this work, we show that during the initial "chaotic" phase of training, even extremely small perturbations reliably causes otherwise identical training trajectories to diverge-an effect that diminishes rapidly over training time. We quantify this divergence through (i) distance between parameters, (ii) the loss barrier when interpolating between networks, (iii) and barrier between parameters after permutation alignment, and (iv) representational similarity between intermediate activations; revealing how perturbations across different hyperparameter or fine-tuning settings drive training trajectories toward distinct loss minima. Our findings provide insights into neural network training stability, with practical implications for fine-tuning, model merging, and diversity of model ensembles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- The Appeal and Reality of Recycling LoRAs with Adaptive MergingHaokun Liu, Gyung Hyun Je, Marco Ciccone, Zhenlin Xu et al.ICML 2026 · 1 citation
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
- Entropic Confinement and Mode Connectivity in Overparameterized Neural NetworksLuca di Carlo, Chase Goddard, David J. SchwabICLR 2026
Builds on33
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Federated Learning with Matched AveragingHongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris S. Papailiopoulos et al.ICLR 2020 · 1,368 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- What is being transferred in transfer learning?Behnam Neyshabur, Hanie Sedghi, Chiyuan ZhangNeurIPS 2020 · 654 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
Related papers
- Chaotic Dynamics are Intrinsic to Neural Network Training with SGDLuis Herrmann, Maximilian Granz, Tim LandgrafNeurIPS 2022 · 15 citations
- Feature-Learning Networks Are Consistent Across Widths At Realistic ScalesNikhil Vyas, Alexander B. Atanasov, Blake Bordelon, Depen Morwani et al.NeurIPS 2023 · 47 citations
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Grounding Representation Similarity Through Statistical TestingFrances Ding, Jean-Stanislas Denain, Jacob SteinhardtNeurIPS 2021 · 88 citations
- On Scrambling Phenomena for Randomly Initialized Recurrent NetworksVaggos Chatziafratis, Ioannis Panageas, Clayton Sanford, Stelios StavroulakisNeurIPS 2022 · 3 citations
