Entropic gradient descent algorithms and wide flat minima
Fabrizio Pittorino, Carlo Lucibello, Christoph Feinauer, Gabriele Perugini, Carlo Baldassi, Elizaveta Demyanenko, Riccardo Zecchina
Abstract
The properties of flat minima in the empirical risk landscape of neural networks have been debated for some time. Increasing evidence suggests they possess better generalization capabilities with respect to sharp ones. In this work we first discuss the relationship between alternative measures of flatness: the local entropy, which is useful for analysis and algorithm development, and the local energy, which is easier to compute and was shown empirically in extensive tests on state-of-the-art networks to be the best predictor of generalization capabilities. We show semi-analytically in simple controlled scenarios that these two measures correlate strongly with each other and with generalization. Then, we extend the analysis to the deep learning scenario by extensive numerical validations. We study two algorithms, entropy-stochastic gradient descent and replicated-stochastic gradient descent, that explicitly include the local entropy in the optimization objective. We devise a training schedule by which we consistently find flatter minima (using both flatness measures), and improve the generalization error for common architectures (e.g. ResNet, EfficientNet).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 379f9c1d-0f81-4fe8-a017-7cf2cff4848fCited by top-tier papers7
- Overcoming Catastrophic Forgetting in Incremental Few-Shot Learning by Finding Flat MinimaGuangyuan Shi, Jiaxin Chen, Wenlong Zhang, Li-Ming Zhan et al.NeurIPS 2021 · 229 citations
- Neural networks with late-phase weightsJohannes von Oswald, Seijin Kobayashi, João Sacramento, Alexander Meulemans et al.ICLR 2021 · 38 citations
- Deep Networks on Toroids: Removing Symmetries Reveals the Structure of Flat Regions in the Landscape GeometryFabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer et al.ICML 2022 · 30 citations
- REPAIR: REnormalizing Permuted Activations for Interpolation RepairKeller Jordan, Hanie Sedghi, Olga Saukh, Rahim Entezari et al.ICLR 2023 · 11 citations
- Entropy-MCMC: Sampling from Flat Basins with EaseBolian Li, Ruqi ZhangICLR 2024 · 7 citations
Builds on1
Related papers
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
- When Do Flat Minima Optimizers Work?Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. KusnerNeurIPS 2022 · 102 citations
- GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved GeneralizationZhiyuan Zhang, Ruixuan Luo, Qi Su, Xu SunEMNLP 2022 · 9 citations
- How to Escape Sharp Minima with Random PerturbationsKwangjun Ahn, Ali Jadbabaie, Suvrit SraICML 2024 · 17 citations
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan et al.NeurIPS 2022 · 132 citations
