Adaptive Preconditioners Trigger Loss Spikes in Adam
Zhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li, Zhiyu li, Feiyu Xiong, Hongkang Yang, Yaoyu Zhang, Zhi-Qin John Xu
摘要
Loss spikes commonly emerge during neural network training with the Adam optimizer across diverse architectures and scales, yet their underlying mechanism remains elusive. While previous explanations attribute these phenomena to sharper loss landscapes at lower loss, we show that landscape geometry alone is insufficient to explain the phenomenon. In this work, we pinpoint the root cause in the internal dynamics of Adam's second moment estimator. We identify a critical ``decoupling'' mechanism where the adaptive preconditioner fails to track the instantaneous squared gradients , causing the adaptive mechanism to effectively fail. This decoupling allows the preconditioner to decay autonomously despite rising gradients, which pushes the maximum eigenvalue of the preconditioned Hessian beyond the stability threshold for sustained periods, manifesting as dramatic loss spikes. Through a quadratic approximation analysis, we theoretically and experimentally characterize five distinct stages of spike evolution and propose a predictor for anticipating spikes based on gradient-directional curvature. We empirically find that the proposed loss spike mechanism, although derived from simplified models, generalizes well to practical scenarios ranging from small neural networks to large-scale Transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient ClippingGuoxia Wang, Shuai Li, Congliang Chen, Jinle Zeng 等ICML 2026 · 被引用 3 次
- Towards Understanding Adam Convergence on Highly Degenerate PolynomialsZhiwei Bai, Jiajie Zhao, Zhangchen Zhou, Zhi-Qin John Xu 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper15
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
- Stabilizing Transformer Training by Preventing Attention Entropy CollapseShuangfei Zhai, Tatiana Likhomanenko, Etai Littwin, Dan Busbridge 等ICML 2023 · 被引用 153 次
- Understanding Gradient Descent on the Edge of Stability in Deep LearningSanjeev Arora, Zhiyuan Li, Abhishek PanigrahiICML 2022 · 被引用 139 次
- Adam Can Converge Without Any Modification On Update RulesYushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun 等NeurIPS 2022 · 被引用 134 次
相关 Paper
- Optimizer Choice Matters For The Emergence of Neural CollapseJim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien LucchiICLR 2026 · 被引用 1 次
- The Expected Loss of Preconditioned Langevin Dynamics Reveals the Hessian RankAmitay Bar, Rotem Mulayoff, Tomer Michaeli, Ronen TalmonAAAI 2024 · 被引用 1 次
- GradientStabilizer: Fix the Norm, Not the GradientTianjin Huang, Zhangyang “Atlas” Wang, Haotian Hu, Zhenyu Zhang 等ICML 2026
- Does SGD really happen in tiny subspaces?Minhak Song, Kwangjun Ahn, Chulhee YunICLR 2025
- STEP: Learning N: M Structured Sparsity Masks from Scratch with PreconditionYucheng Lu, Shivani Agrawal, Suvinay Subramanian, Oleg Rybakov 等ICML 2023 · 被引用 32 次
