When Do Flat Minima Optimizers Work?
Jean Kaddour, Linqing Liu, Ricardo Silva, Matt J. Kusner
Abstract
Recently, flat-minima optimizers, which seek to find parameters in low-loss neighborhoods, have been shown to improve a neural network's generalization performance over stochastic and adaptive gradient-based optimizers. Two methods have received significant attention due to their scalability: 1. Stochastic Weight Averaging (SWA), and 2. Sharpness-Aware Minimization (SAM). However, there has been limited investigation into their properties and no systematic benchmarking of them across different domains. We fill this gap here by comparing the loss surfaces of the models trained with each method and through broad benchmarking across computer vision, natural language processing, and graph representation learning tasks. We discover several surprising findings from these results, which we hope will help researchers further improve deep learning optimizers, and practitioners identify the right optimizer for their problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f655a703-7140-4d13-bfd8-fc673ba57042Cited by top-tier papers42
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy et al.NeurIPS 2022 · 183 citations
- No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsJean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini et al.NeurIPS 2023 · 63 citations
- Mechanistic Mode ConnectivityEkdeep Singh Lubana, Eric J. Bigelow, Robert P. Dick, David Scott Krueger et al.ICML 2023 · 57 citations
- PLASTIC: Improving Input and Label Plasticity for Sample Efficient Reinforcement LearningHojoon Lee, Hanseul Cho, Hyunseung Kim, Daehoon Gwak et al.NeurIPS 2023 · 50 citations
- Normalization Layers Are All That Sharpness-Aware Minimization NeedsMaximilian Müller, Tiffany Vlaar, David Rolnick, Matthias HeinNeurIPS 2023 · 37 citations
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Sharpness-Aware Minimization Revisited: Weighted Sharpness as a Regularization TermYun Yue, Jiadi Jiang, Zhiling Ye, Ning Gao et al.KDD 2023 · 7 citations
- Stability Analysis of Sharpness-Aware MinimizationHoki Kim, Jinseong Park, Yujin Choi, Jaewook LeeICML 2026 · 18 citations
- GA-SAM: Gradient-Strength based Adaptive Sharpness-Aware Minimization for Improved GeneralizationZhiyuan Zhang, Ruixuan Luo, Qi Su, Xu SunEMNLP 2022 · 9 citations
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou et al.CVPR 2023
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
