Lune

ICML2025顶会

Avoiding spurious sharpness minimization broadens applicability of SAM

Sidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. Dauphin

出版方
2025年份
6顶会引用

摘要

Curvature regularization techniques like Sharpness Aware Minimization (SA M) have shown great promise in improving generalization on vision tasks. However, we find that SA M performs poorly in domains like natural language processing (NLP), often degrading performance -even with twice the compute budget. We investigate the discrepancy across domains and find that in the NLP setting, SAM is dominated by regularization of the logit statistics --instead of improving the geometry of the function itself. We use this observation to develop an alternative algorithm we call Fu nc t i onal -SA M, which regularizes curvature only through modification of the statistics of the overall function implemented by the neural network, and avoids spurious minimization through logit manipulation. Furthermore, we argue that preconditioning the SA M perturbation also prevents spurious minimization, and when combined with Fu nc t i onal -SA M, it gives further improvements. Our proposed algorithms show improved performance over A dam W and SA M baselines when trained for an equal number of steps, in both fixed-length and Chinchilla-style training settings, at various model scales (including billionparameter scale). On the whole, our work highlights the importance of more precise characterizations of sharpness in broadening the applicability of curvature regularization to large language models (LLMs).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper26

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖