Nondeterminism and Instability in Neural Network Optimization
Cecilia Summers, Michael J. Dinneen
摘要
Nondeterminism in neural network optimization produces uncertainty in performance, making small improvements difficult to discern from runto-run variability. While uncertainty can be reduced by training multiple model copies, doing so is time-consuming, costly, and harms reproducibility. In this work, we establish an experimental protocol for understanding the effect of optimization nondeterminism on model diversity, allowing us to isolate the effects of a variety of sources of nondeterminism. Surprisingly, we find that all sources of nondeterminism have similar effects on measures of model diversity. To explain this intriguing fact, we identify the instability of model training, taken as an end-to-end procedure, as the key determinant. We show that even onebit changes in initial parameters result in models converging to vastly different values. Last, we propose two approaches for reducing the effects of instability on run-to-run variability. Introduction Consider this common scenario: you have a baseline "current best" model, and are trying to improve it. One of your experiments has produced a model whose metrics are slightly better than the baseline. Yet you have your reservations -how do you know the improvement is "real" and not due to run-to-run variability? Similarly, consider hyperparameter optimization, in which many possible values exist for a set of hyperparameters, with minor differences in performance between them. How do you pick the best hyperparameters, and how can you be sure that you've actually picked wisely? In both scenarios, the standard practice is to train multiple independent copies of your model to understand its variability. While this helps address the problem, it is extremely
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Elo Uncovered: Robustness and Best Practices in Language Model EvaluationMeriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker 等NeurIPS 2024 · 被引用 94 次
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM InferenceJiayi Yuan, Hao Li, Xinheng Ding, Wenya Xie 等NeurIPS 2025 · 被引用 74 次
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 被引用 25 次
- On the Variance of Neural Network Training with respect to Test Sets and DistributionsKeller JordanICLR 2024 · 被引用 23 次
- On The Fairness Impacts of Hardware Selection in Machine LearningSree Harsha Nelaturu, Nishaanth Kanna Ravichandran, Cuong Tran, Sara Hooker 等ICML 2024 · 被引用 5 次
它引用的顶会 Paper2
相关 Paper
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial ConditionsGül Sena Altintas, Devin Kwok, Colin Raffel, David RolnickICML 2025
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier 等ASE 2020 · 被引用 91 次
- Causes and Effects of Unanticipated Numerical Deviations in Neural Network Inference FrameworksAlexander Schlögl, Nora Hofer, Rainer BöhmeNeurIPS 2023 · 被引用 31 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
- On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsRuoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi 等ACL 2026
