Improving SAM Requires Rethinking its Optimization Formulation
Wanyun Xie, Fabian Latorre, Kimon Antonakopoulos, Thomas Pethick, Volkan Cevher
Abstract
This paper rethinks Sharpness-Aware Minimization (SAM), which is originally formulated as a zero-sum game where the weights of a network and a bounded perturbation try to minimize/maximize, respectively, the same differentiable loss. To fundamentally improve this design, we argue that SAM should instead be reformulated using the 0-1 loss. As a continuous relaxation, we follow the simple conventional approach where the minimizing (maximizing) player uses an upper bound (lower bound) surrogate to the 0-1 loss. This leads to a novel formulation of SAM as a bilevel optimization problem, dubbed as BiSAM. BiSAM with newly designed lower-bound surrogate loss indeed constructs stronger perturbation. Through numerical evidence, we show that BiSAM consistently results in improved performance when compared to the original SAM and variants, while enjoying similar computational complexity. Our code is available at https://github.com/ LIONS-EPFL/BiSAM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 400e6068-eb84-4c19-bb34-8377cec1a95dCited by top-tier papers5
- Bilevel Optimization for Adversarial Learning Problems: Sharpness, Generation, and BeyondRisheng Liu, Zhu Liu, Weihao Mao, Wei Yao et al.NeurIPS 2025 · 2 citations
- Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded SchedulerDimitris Oikonomou, Nicolas LoizouICML 2026
- Tilted Sharpness-Aware MinimizationTian Li, Tianyi Zhou, Jeff A. BilmesICML 2025
- Understanding SAM through Minimax PerspectiveYing Chen, Aoxi Li, Javad LavaeiICML 2026
- Flatness-Aware Stochastic Gradient Langevin DynamicsStefano Bruno, Youngsik Hwang, JaeHyeon An, Sotirios Sabanis et al.ICML 2026
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui et al.ICLR 2022 · 213 citations
Related papers
- Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware MinimizationJinping Wang, Qinhan Liu, Zhiwu Xie, Zhiqiang GaoICML 2026
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 1 citation
- Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves GeneralizationXingxuan Zhang, Renzhe Xu, Han Yu, Hao Zou et al.CVPR 2023
- Towards Understanding Sharpness-Aware MinimizationMaksym Andriushchenko, Nicolas FlammarionICML 2022 · 190 citations
- Why is SAM Robust to Label Noise?Christina Baek, J. Zico Kolter, Aditi RaghunathanICLR 2024 · 24 citations
