Towards Efficient and Scalable Sharpness-Aware Minimization
Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, Yang You
Abstract
Recently, Sharpness-Aware Minimization (SAM), which connects the geometry of the loss landscape and generalization, has demonstrated a significant performance boost on training large-scale models such as vision transformers. However, the update rule of SAM requires two sequential (non-parallelizable) gradient computations at each step, which can double the computational overhead. In this paper, we propose a novel algorithm LookSAM - that only periodically calculates the inner gradient ascent, to significantly reduce the additional training cost of SAM. The empirical results illustrate that LookSAM achieves similar accuracy gains to SAM while being tremendously faster - it enjoys comparable computational complexity with first-order optimizers such as SGD or Adam. To further evaluate the performance and scalability of LookSAM, we incorporate a layer-wise modification and perform experiments in the large-batch training scenario, which is more prone to converge to sharp local minima. Equipped with the proposed algorithms, we are the first to successfully scale up the batch size when training Vision Transformers (ViTs). With a 64k batch size, we are able to train ViTs from scratch in minutes while maintaining competitive performance. The code is available here: https://github.com/yong-6/LookSAM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52dcea80-d55e-411e-a4bc-6612504bdbd0Cited by top-tier papers84
- Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsZhiyuan Liu, Hong Liu, Denny Zhou, Tengyu MaICLR 2024 · 259 citations
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou et al.ICLR 2022 · 168 citations
- Sharpness-Aware Training for FreeJiawei Du, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan et al.NeurIPS 2022 · 132 citations
- Improving the Model Consistency of Decentralized Federated LearningYifan Shi, Li Shen, Kang Wei, Yan Sun et al.ICML 2023 · 89 citations
- Escaping Saddle Points for Effective Generalization on Class-Imbalanced DataHarsh Rangwani, Sumukh K. Aithal, Mayank Mishra, Venkatesh Babu R.NeurIPS 2022 · 50 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data AugmentationsXiangning Chen, Cho-Jui Hsieh, Boqing GongICLR 2022 · 388 citations
Related papers
- Random Sharpness-Aware MinimizationYong Liu, Siqi Mai, Minhao Cheng, Xiangning Chen et al.NeurIPS 2022 · 38 citations
- Efficient Sharpness-Aware Minimization for Molecular Graph Transformer ModelsYili Wang, Kaixiong Zhou, Ninghao Liu, Ying Wang et al.ICLR 2024 · 13 citations
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui et al.ICLR 2022 · 213 citations
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 16 citations
- SAMPa: Sharpness-aware Minimization ParallelizedWanyun Xie, Thomas Pethick, Volkan CevherNeurIPS 2024 · 12 citations
