On the Training Instability of Shuffling SGD with Batch Normalization
David Xing Wu, Chulhee Yun, Suvrit Sra
摘要
We uncover how SGD interacts with batch normalization and can exhibit undesirable training dynamics such as divergence. More precisely, we study how Single Shuffle (SS) and Random Reshuffle (RR) -- two widely used variants of SGD -- interact surprisingly differently in the presence of batch normalization: RR leads to much more stable evolution of training loss than SS. As a concrete example, for regression using a linear network with batch normalization, we prove that SS and RR converge to distinct global optima that are"distorted"away from gradient descent. Thereafter, for classification we characterize conditions under which training divergence for SS and RR can, and cannot occur. We present explicit constructions to show how SS leads to distorted optima in regression and divergence for classification, whereas RR avoids both distortion and divergence. We validate our results by confirming them empirically in realistic settings, and conclude that the separation between SS and RR used with batch normalization is relevant in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bounded and Unbiased Composite Differential PrivacyKai Zhang, Yanjun Zhang, Ruoxi Sun, Pei-Wei Tsai 等S&P 2024 · 被引用 54 次
- Diffusion On Syntax Trees For Program SynthesisShreyas Kapur, Erik Jenner, Stuart RussellICLR 2025
它引用的顶会 Paper14
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 被引用 267 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
相关 Paper
- On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight DecayEkaterina Lobacheva, Maxim Kodryan, Nadezhda Chirkova, Andrey Malinin 等NeurIPS 2021 · 被引用 30 次
- Understanding the Impact of Model Incoherence on Convergence of Incremental SGD with Random ReshuffleShaocong Ma, Yi ZhouICML 2020 · 被引用 5 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- A Theoretical Analysis of the Learning Dynamics under Class ImbalanceEmanuele Francazi, Marco Baity-Jesi, Aurélien LucchiICML 2023 · 被引用 33 次
- Momentum Further Constrains Sharpness at the Edge of Stochastic StabilityArseniy Andreyev, Advikar Ananthkumar, Marc Walden, Tomaso A Poggio 等ICML 2026 · 被引用 5 次
