A Better Start: Sensitivity-Aware Warm-Up for Robust and Efficient Fine-Tuning
Yile Chen, Zeyi Wen, Jian Chen, Jin Huang
摘要
As an essential component of fine-tuning, warm-up plays a crucial role in promoting stability and generalization. Many studies have examined its underlying mechanisms from different aspects. However, most of the studies focus on incorporating these insights into optimizers to reduce the reliance on warm-up. Little attention has been paid to addressing the inherent limitations of the warm-up itself, which restricts its effectiveness. In this work, we revisit warm-up from a loss landscape perspective and identify several limitations with existing warm-up, including: (1) susceptibility to nearby suboptimal traps, (2) sensitivity to hyperparameters and random seeds, and (3) inefficiency during the early stages of training. To overcome these limitations, we propose Sensitivity-Aware Warm-Up (SAWU), a lightweight and adaptive strategy that dynamically leverages learning sensitivity during warm-up to guide updates toward better and more stable basins. In addition, SAWU also introduces an adaptive scheduling mechanism and phase transition strategy across warm-up, stable, and decay phases to further enhance robustness and efficiency. Extensive experiments on various downstream tasks show that SAWU significantly outperforms the vanilla method (e.g., average 3.43% improvement on RoBerta). Moreover, SAWU can be easily combined with various optimizers and remains effective even when warm-up-based methods fail (e.g, it lifts RAdam from 49.46% to 91.78% on qnli. Thanks to its lightweight nature, SAWU introduces minimal overhead and even reduces training time by over 5% compared to other methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Why Warmup the Learning Rate? Underlying Mechanisms and ImprovementsDayal Singh Kalra, Maissam BarkeshliNeurIPS 2024 · 被引用 87 次
相关 Paper
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase 等ICLR 2026 · 被引用 13 次
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li 等EMNLP 2022 · 被引用 5 次
- Sensitivity-Aware Visual Parameter-Efficient Fine-TuningHaoyu He, Jianfei Cai, Jing Zhang, Dacheng Tao 等ICCV 2023 · 被引用 97 次
- Beyond Static Allocation: Dynamic Sensitivity-Aware Fine-Tuning for Vision TransformersYuanyang Cao, Xichun Liu, Fuwei Zhang, Shangqi Deng 等ICML 2026
- No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer ModelsChen Liang, Haoming Jiang, Simiao Zuo, Pengcheng He 等ICLR 2022 · 被引用 18 次
