Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant Problems
Bingcong Li, Liang Zhang, Niao He
Abstract
Sharpness-aware minimization (SAM) improves generalization of various deep learning tasks. Motivated by popular architectures such as LoRA, we explore the implicit regularization of SAM for scale-invariant problems involving two groups of variables. Instead of focusing on commonly used sharpness, this work introduces a concept termed balancedness, defined as the difference between the squared norm of two variables. This allows us to depict richer global behaviors of SAM. In particular, our theoretical and empirical findings reveal that i) SAM promotes balancedness; and ii) the regularization on balancedness is data-responsive -- outliers have stronger impact. The latter coincides with empirical observations that SAM outperforms SGD in the presence of outliers. Leveraging the implicit regularization, we develop a resource-efficient SAM variant, balancedness-aware regularization (BAR), tailored for scale-invariant problems such as finetuning language models with LoRA. BAR saves 95% computational overhead of SAM, with enhanced test performance across various tasks on RoBERTa, GPT2, and OPT-1.3B.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Stability Analysis of Sharpness-Aware MinimizationHoki Kim, Jinseong Park, Yujin Choi, Jaewook LeeICML 2026 · 18 citations
- RefLoRA: Refactored Low-Rank Adaptation for Efficient Fine-Tuning of Large ModelsYilang Zhang, Bingcong Li, Georgios B. GiannakisNeurIPS 2025 · 9 citations
- Zeroth-Order Optimization Finds Flat MinimaLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh et al.NeurIPS 2025 · 8 citations
- Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale ModelsYuhang Liu, Tao Li, Zhehao Huang, Zuopeng Yang et al.ICLR 2026 · 3 citations
- Efficiently Seeking Flat Minima for Better Generalization in Fine-Tuning Large Language Models and BeyondJiaxin Deng, Qingcheng Zhu, Junbiao Pang, Linlin Yang et al.AAAI 2026 · 1 citation
Builds on34
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
Related papers
- Balanced LoRA: Removing Parameter Invariance to Accelerate ConvergenceValérie Castin, Kimia Nadjahi, Pierre Ablin, Gabriel PeyréICML 2026
- Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized ModelsTianxiao Cao, Kyohei Atarashi, Hisashi KashimaAAAI 2026
- How Sharpness-Aware Minimization Minimizes Sharpness?Kaiyue Wen, Tengyu Ma, Zhiyuan LiICLR 2023 · 3 citations
- IsoBN: Fine-Tuning BERT with Isotropic Batch NormalizationWenxuan Zhou, Bill Yuchen Lin, Xiang RenAAAI 2021 · 29 citations
- BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language ModelsYupeng Chang, Yi Chang, Yuan WuICLR 2026 · 3 citations
