Nondeterminism and Instability in Neural Network Optimization
Cecilia Summers, Michael J. Dinneen
Abstract
Nondeterminism in neural network optimization produces uncertainty in performance, making small improvements difficult to discern from runto-run variability. While uncertainty can be reduced by training multiple model copies, doing so is time-consuming, costly, and harms reproducibility. In this work, we establish an experimental protocol for understanding the effect of optimization nondeterminism on model diversity, allowing us to isolate the effects of a variety of sources of nondeterminism. Surprisingly, we find that all sources of nondeterminism have similar effects on measures of model diversity. To explain this intriguing fact, we identify the instability of model training, taken as an end-to-end procedure, as the key determinant. We show that even onebit changes in initial parameters result in models converging to vastly different values. Last, we propose two approaches for reducing the effects of instability on run-to-run variability. Introduction Consider this common scenario: you have a baseline "current best" model, and are trying to improve it. One of your experiments has produced a model whose metrics are slightly better than the baseline. Yet you have your reservations -how do you know the improvement is "real" and not due to run-to-run variability? Similarly, consider hyperparameter optimization, in which many possible values exist for a set of hyperparameters, with minor differences in performance between them. How do you pick the best hyperparameters, and how can you be sure that you've actually picked wisely? In both scenarios, the standard practice is to train multiple independent copies of your model to understand its variability. While this helps address the problem, it is extremely
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 29d9624c-5cea-4fae-a458-b9c5fdf8c107Cited by top-tier papers12
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker et al.NeurIPS 2024 · 94 citations
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM InferenceJiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie et al.NeurIPS 2025 · 74 citations
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 25 citations
- On the Variance of Neural Network Training with respect to Test Sets and DistributionsKeller JordanICLR 2024 · 23 citations
- On The Fairness Impacts of Hardware Selection in Machine LearningSree Harsha Nelaturu, Nishaanth Kanna Ravichandran, Cuong Tran, Sara Hooker et al.ICML 2024 · 5 citations
Builds on2
Related papers
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial ConditionsGül Sena Altintas, Devin Kwok, Colin Raffel, David RolnickICML 2025
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 31 citations
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett et al.ICLR 2024 · 162 citations
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi et al.ACL 2026
