Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
Liu Ziyin, Mingze Wang, Hongchao Li, Lei Wu
Abstract
Symmetries are prevalent in deep learning and can significantly influence the learning dynamics of neural networks. In this paper, we examine how exponential symmetries -- a broad subclass of continuous symmetries present in the model architecture or loss function -- interplay with stochastic gradient descent (SGD). We first prove that gradient noise creates a systematic motion (a ``Noether flow") of the parameters along the degenerate direction to a unique initialization-independent fixed point . These points are referred to as the noise equilibria because, at these points, noise contributions from different directions are balanced and aligned. Then, we show that the balance and alignment of gradient noise can serve as a novel alternative mechanism for explaining important phenomena such as progressive sharpening/flattening and representation formation within neural networks and have practical implications for understanding techniques like representation normalization and warmup.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Neural Thermodynamics: Entropic Forces in Deep and Universal Representation LearningLiu Ziyin, Yizhou Xu, Isaac L. ChuangNeurIPS 2025 · 11 citations
- Differentiable Sparsity via -Gating: Simple and Versatile Structured PenalizationChris Kolb, Laetitia Frost, Bernd Bischl, David RügamerNeurIPS 2025 · 4 citations
- Formation of Representations in Neural NetworksLiu Ziyin, Isaac L. Chuang, Tomer Galanti, Tomaso A. PoggioICLR 2025 · 1 citation
- Remove Symmetries to Control Model Expressivity and Improve OptimizationLiu Ziyin, Yizhou Xu, Isaac L. ChuangICLR 2025
- Beyond the Permutation Symmetry of Transformers: The Role of Rotation for Model FusionBinchi Zhang, Zaiyi Zheng, Zhengzhang Chen, Jundong LiICML 2025
Builds on18
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 135 citations
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)Zhiyuan Li, Sadhika Malladi, Sanjeev AroraNeurIPS 2021 · 107 citations
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins et al.ICLR 2021 · 100 citations
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning RateZhiyuan Li, Kaifeng Lyu, Sanjeev AroraNeurIPS 2020 · 93 citations
Related papers
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural NetworksHidenori Tanaka, Daniel KuninNeurIPS 2021 · 57 citations
- Symmetry Induces Structure and Constraint of LearningLiu ZiyinICML 2024 · 24 citations
- Symmetries in Overparametrized Neural Networks: A Mean Field ViewJavier Maass Martínez, Joaquín FontbonaNeurIPS 2024 · 4 citations
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 52 citations
- SGD with Large Step Sizes Learns Sparse FeaturesMaksym Andriushchenko, Aditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionICML 2023 · 77 citations
