Exploiting weight-space symmetries for approximating curvature
Artem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu, Felix Dangel, Guillaume Hennequin, Alberto Bernacchia
Abstract
Many machine learning techniques rely on approximating a loss function's curvature, but this is notoriously hard to do at the scale of modern deep networks. Surprisingly, no previous work has exploited the curvature constraints that arise from well known weight-space symmetries in loss landscapes. By analytically averaging over group actions that leave the loss invariant, we construct structured Hessian approximations from single gradients that can be tractably estimated, stored, and inverted. The choice of user-specified symmetry group directly governs the trade-off between approximation accuracy and computational cost. Moreover, our framework provides a unifying theoretical lens for viewing existing methods; in particular, a specific choice of symmetry group recovers Shampoo/Muon-like curvature estimates. We validate our method on a range of network architectures, and deploy it to second-order optimization benchmarks, including a small language model. Our curvature estimation framework might find applications in other machine learning problems such as uncertainty estimation, continual learning, compression/pruning, training data attribution, and more.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0b3ae4a-eb60-43de-80e0-a153e6bca62aBuilds on23
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 301 citations
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and InvariancesBerfin Simsek, François Ged, Arthur Jacot, Francesco Spadaro et al.ICML 2021 · 136 citations
- Practical Quasi-Newton Methods for Training Deep Neural NetworksDonald Goldfarb, Yi Ren, Achraf BahamouNeurIPS 2020 · 130 citations
- The Polar Express: Optimal Matrix Sign Methods and their Application to the Muon AlgorithmNoah Amsel, David Persson, Christopher Musco, Robert M. GowerICLR 2026 · 115 citations
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins et al.ICLR 2021 · 100 citations
Related papers
- Global curvature for second-order optimization of neural networksAlberto BernacchiaICML 2025
- M-FAC: Efficient Matrix-Free Approximations of Second-Order InformationElias Frantar, Eldar Kurtic, Dan AlistarhNeurIPS 2021 · 69 citations
- Symmetries, Flat Minima, and the Conserved Quantities of Gradient FlowBo Zhao, Iordan Ganev, Robin Walters, Rose Yu et al.ICLR 2023 · 2 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
- Debiasing Mini-Batch Quadratics for Applications in Deep LearningLukas Tatzel, Bálint Mucsányi, Osane Hackel, Philipp HennigICLR 2025
