The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
Eitan Gronich, Gal Vardi
摘要
We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that momentum steepest descent algorithms like Muon (spectral norm), MomentumGD ( norm), and Signum ( norm) are approximate steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong 等NeurIPS 2020 · 被引用 309 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
- Adam Can Converge Without Any Modification On Update RulesYushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun 等NeurIPS 2022 · 被引用 134 次
相关 Paper
- Implicit Bias of Spectal Descent and Muon on Multiclass Separable DataChen Fan, Mark Schmidt, Christos ThrampoulidisNeurIPS 2025
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 被引用 1 次
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many VariantsMichael Crawshaw, Chirag Modi, Mingrui Liu, Robert GowerICML 2026 · 被引用 24 次
- Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural NetworksNikolaos Tsilivis, Gal Vardi, Julia KempeICLR 2025
- Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch RegimeBeomhan Baek, Minhak Song, Chulhee YunICLR 2026 · 被引用 2 次
