Never Saddle for Reparameterized Steepest Descent as Mirror Flow
Tom Jacobs, Chao Zhou, Rebekka Burkholz
摘要
How does the choice of optimization algorithm shape a model’s ability to learn features? To address this question for steepest descent methods --including sign descent, which is closely related to Adam --we introduce steepest mirror flows as a unifying theoretical framework. This framework reveals how optimization geometry governs learning dynamics, implicit bias, and sparsity and it provides two explanations for why Adam and AdamW often outperform SGD in fine-tuning. Focusing on diagonal linear networks and deep diagonal linear reparameterizations (a simplified proxy for attention), we show that steeper descent facilitates both saddle-point escape and feature learning. In contrast, gradient descent requires unrealistically large learning rates to escape saddles, an uncommon regime in fine-tuning. Empirically, we confirm that saddle-point escape is a central challenge in fine-tuning. Furthermore, we demonstrate that decoupled weight decay, as in AdamW, stabilizes feature learning by enforcing novel balance equations. Together, these results highlight two mechanisms how steepest descent can aid modern optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 135 次
- Robustness to Unbounded Smoothness of Generalized SignSGDMichael Crawshaw, Mingrui Liu, Francesco Orabona, Wei Zhang 等NeurIPS 2022 · 被引用 111 次
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 被引用 92 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
相关 Paper
- Implicit Bias of AdamW: ℓ∞-Norm Constrained OptimizationShuo Xie, Zhiyuan LiICML 2024 · 被引用 46 次
- Optimizer Choice Matters For The Emergence of Neural CollapseJim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien LucchiICLR 2026 · 被引用 1 次
- Stacey: Promoting Stochastic Steepest Descent via Accelerated ℓp-Smooth Nonconvex OptimizationXinyu Luo, Site Bai, Bolian Li, Petros Drineas 等ICML 2025
- Cautious Weight DecayLizhang Chen, Jonathan Li, Kaizhao Liang, Baiyu Su 等ICLR 2026 · 被引用 14 次
- Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural NetworksNikolaos Tsilivis, Gal Vardi, Julia KempeICLR 2025
