On the Training Instability of Shuffling SGD with Batch Normalization
David Xing Wu, Chulhee Yun, Suvrit Sra
Abstract
We uncover how SGD interacts with batch normalization and can exhibit undesirable training dynamics such as divergence. More precisely, we study how Single Shuffle (SS) and Random Reshuffle (RR) -- two widely used variants of SGD -- interact surprisingly differently in the presence of batch normalization: RR leads to much more stable evolution of training loss than SS. As a concrete example, for regression using a linear network with batch normalization, we prove that SS and RR converge to distinct global optima that are"distorted"away from gradient descent. Thereafter, for classification we characterize conditions under which training divergence for SS and RR can, and cannot occur. We present explicit constructions to show how SS leads to distorted optima in regression and divergence for classification, whereas RR avoids both distortion and divergence. We validate our results by confirming them empirically in realistic settings, and conclude that the separation between SS and RR used with batch normalization is relevant in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e25206fa-77ee-4614-a088-0cc1e02b16d0Cited by top-tier papers2
- Bounded and Unbiased Composite Differential PrivacyKai Zhang, Yanjun Zhang, Ruoxi Sun, Pei-Wei Tsai et al.S&P 2024 · 54 citations
- Diffusion On Syntax Trees For Program SynthesisShreyas Kapur, Erik Jenner, Stuart RussellICLR 2025
Builds on14
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 111 citations
Related papers
- On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight DecayEkaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin et al.NeurIPS 2021 · 30 citations
- Understanding the Impact of Model Incoherence on Convergence of Incremental SGD with Random ReshuffleShaocong Ma, Yi ZhouICML 2020 · 5 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- A Theoretical Analysis of the Learning Dynamics under Class ImbalanceEmanuele Francazi, Marco Baity-Jesi, Aurélien LucchiICML 2023 · 33 citations
- Momentum Further Constrains Sharpness at the Edge of Stochastic StabilityArseniy Andreyev, Advikar Ananthkumar, Marc Walden, Tomaso A Poggio et al.ICML 2026 · 5 citations
