Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
Ishaan Watts, Catherine Li, Sachin Goyal, Jacob Mitchell Springer, Aditi Raghunathan
摘要
Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger model after subsequent changes like post-training and quantization. This overlooks the geometry of the base model which controls how much of the base model's capabilities survive subsequent parameter updates. We study three pretraining optimization approaches that bias optimization toward flatter minima: Sharpness-Aware Minimization (SAM), large learning rates, and shortened learning rate annealing periods. Across model sizes ranging from 20M to 150M parameters, we find that these interventions consistently improve downstream performance after post-training on five common datasets with up to 80% less forgetting. These principles hold at scale: a short SAM mid-training phase applied to an existing OLMo-2-1B checkpoint reduces forgetting by 31% after MetaMath post-training and by 40% after 4-bit quantization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural NetworksJungmin Kwon, Jeongseop Kim, Hyunseo Park, In Kwon ChoiICML 2021 · 被引用 385 次
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 被引用 288 次
相关 Paper
- Beyond Outliers: A Study of Optimizers Under QuantizationGeorgios Vlassis, Saleh Ashkboos, Alexandra Volkova, Torsten Hoefler 等ICLR 2026 · 被引用 6 次
- Asymptotic Unbiased Sample Sampling to Speed Up Sharpness-Aware MinimizationJiaxin Deng, Junbiao Pang, Baochang Zhang, Guodong GuoAAAI 2025 · 被引用 5 次
- Forget Sharpness: Perturbed Forgetting of Model Biases Within SAM DynamicsAnkit Vani, Frederick Tung, Gabriel L. Oliveira, Hossein Sharifi-NoghabiICML 2024
- Upweighting Easy Samples in Fine-Tuning Mitigates ForgettingSunny Sanyal, Hayden Prairie, Rudrajit Das, Ali Kavis 等ICML 2025
- Mapping Post-Training Forgetting in Language Models at ScaleJackson Harmon, Andreas Hochlehnert, Matthias Bethge, Ameya PrabhuICLR 2026 · 被引用 8 次
