Symmetries, Flat Minima, and the Conserved Quantities of Gradient Flow
Bo Zhao, Iordan Ganev, Robin Walters, Rose Yu, Nima Dehmamy
Abstract
Empirical studies of the loss landscape of deep networks have revealed that many local minima are connected through low-loss valleys. Yet, little is known about the theoretical origin of such valleys. We present a general framework for finding continuous symmetries in the parameter space, which carve out low-loss valleys. Our framework uses equivariances of the activation functions and can be applied to different layer architectures. To generalize this framework to nonlinear neural networks, we introduce a novel set of nonlinear, data-dependent symmetries. These symmetries can transform a trained model such that it performs similarly on new samples, which allows ensemble building that improves robustness under certain adversarial attacks. We then show that conserved quantities associated with linear symmetries can be used to define coordinates along low-loss valleys. The conserved quantities help reveal that using common initialization methods, gradient flow only explores a small part of the global minimum. By relating conserved quantities to convergence rate and sharpness of the minimum, we provide insights on how initialization impacts convergence and generalizability. * Equal contribution. † Work done during an internship at IBM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60a86eb7-acbe-452b-a9b5-87003b65db11Cited by top-tier papers16
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 54 citations
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Improving Convergence and Generalization Using Parameter SymmetriesBo Zhao, Robert M. Gower, Robin Walters, Rose YuICLR 2024 · 24 citations
- Class Incremental Learning with Multi-Teacher DistillationHaitao Wen, Lili Pan, Yu Dai, Heqian Qiu et al.CVPR 2024 · 22 citations
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 15 citations
Builds on11
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
- Relative Flatness and GeneralizationHenning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu et al.NeurIPS 2021 · 114 citations
- Loss Surface Simplexes for Mode Connecting Volumes and Fast EnsemblingGregory W. Benton, Wesley J. Maddox, Sanae Lotfi, Andrew Gordon WilsonICML 2021 · 88 citations
- Understanding the Dynamics of Gradient Flow in Overparameterized Linear modelsSalma Tarmoun, Guilherme França, Benjamin D. Haeffele, René VidalICML 2021 · 76 citations
Related papers
- A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant ModelsYuqing Xie, Tess E. SmidtNeurIPS 2025 · 9 citations
- Exploiting weight-space symmetries for approximating curvatureArtem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu et al.ICML 2026
- Learning Layer-wise Equivariances Automatically using GradientsTycho F. A. van der Ouderaa, Alexander Immer, Mark van der WilkNeurIPS 2023 · 28 citations
- Piecewise linear activations substantially shape the loss surfaces of neural networksFengxiang He, Bohan Wang, Dacheng TaoICLR 2020 · 33 citations
- Optimizing Mode Connectivity via Neuron AlignmentN. Joseph Tatro, Pin-Yu Chen, Payel Das, Igor Melnyk et al.NeurIPS 2020 · 104 citations
