Lookbehind-SAM: k steps back, 1 step forward
Gonçalo Mordido, Pranshu Malviya, Aristide Baratin, Sarath Chandar
Abstract
Sharpness-aware minimization (SAM) methods have gained increasing popularity by formulating the problem of minimizing both loss value and loss sharpness as a minimax objective. In this work, we increase the efficiency of the maximization and minimization parts of SAM's objective to achieve a better loss-sharpness tradeoff. By taking inspiration from the Lookahead optimizer, which uses multiple descent steps ahead, we propose Lookbehind, which performs multiple ascent steps behind to enhance the maximization step of SAM and find a worst-case perturbation with higher loss. Then, to mitigate the variance in the descent step arising from the gathered gradients across the multiple ascent steps, we employ linear interpolation to refine the minimization step. Lookbehind leads to a myriad of benefits across a variety of tasks. Particularly, we show increased generalization performance, greater robustness against noisy weights, as well as improved learning and less catastrophic forgetting in lifelong learning settings. Our code is available at https://github. com/chandar-lab/Lookbehind-SAM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseHaocheng Luo, Mehrtash Harandi, Dinh Phung, Trung LeNeurIPS 2025 · 2 citations
- Revisiting Sharpness-Aware Minimization: A More Faithful and Effective ImplementationJianlong Chen, Zhiming ZhouICLR 2026 · 1 citation
- Align-SAM: Seeking Flatter Minima for Better Cross-Subset AlignmentVan-Anh Nguyen, Mehrtash Harandi, Thanh-Toan Do, Linh Ngo Van et al.ICLR 2026
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 385 citations
- Surrogate Gap Minimization Improves Sharpness-Aware TrainingJuntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui et al.ICLR 2022 · 213 citations
- Self-supervised Learning is More Robust to Dataset ImbalanceHong Liu, Jeff Z. HaoChen, Adrien Gaidon, Tengyu MaICLR 2022 · 190 citations
Related papers
- Improving Sharpness-Aware Minimization by LookaheadRunsheng Yu, Youzhi Zhang, James T. KwokICML 2024 · 1 citation
- Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual LearningYanan Chen, Tieliang Gong, Yunjiao Zhang, Wen WenAAAI 2026
- Friendly Sharpness-Aware MinimizationTao Li, Pan Zhou, Zhengbao He, Xinwen Cheng et al.CVPR 2024
- SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local PerturbationHao Ban, Gokul Ram Subramani, Kaiyi JiICCV 2025 · 3 citations
- Momentum-SAM: Sharpness Aware Minimization without Computational OverheadMarlon Becker, Frederick Altrock, Benjamin RisseNeurIPS 2025 · 16 citations
