Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
Nikolaos Tsilivis, Gal Vardi, Julia Kempe
2025年份
11顶会引用
摘要
We study the implicit bias of the general family of steepest descent algorithms with infinitesimal learning rate in deep homogeneous neural networks. We show that: (a) an algorithm-dependent geometric margin starts increasing once the networks reach perfect training accuracy, and (b) any limit point of the training trajectory corresponds to a KKT point of the corresponding margin-maximization problem. We experimentally zoom into the trajectories of neural networks optimized with various steepest descent algorithms, highlighting connections to the implicit bias of popular adaptive methods (Adam and Shampoo).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many VariantsMichael Crawshaw, Chirag Modi, Mingrui Liu, Robert GowerICML 2026 · 被引用 24 次
- How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced DataBhavya Vasudeva, Puneesh Deora, Yize Zhao, Vatsal Sharan 等ICLR 2026 · 被引用 16 次
- The Rich and the Simple: On the Implicit Bias of Adam and SGDBhavya Vasudeva, Jung Hoon Lee, Vatsal Sharan, Mahdi SoltanolkotabiNeurIPS 2025 · 被引用 14 次
- Never Saddle for Reparameterized Steepest Descent as Mirror FlowTom Jacobs, Chao Zhou, Rebekka BurkholzICLR 2026 · 被引用 3 次
- Any-stepsize Gradient Descent for Separable Data under Fenchel-Young LossesHan Bao, Shinsaku Sakaue, Yuki TakezawaNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper11
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir 等NeurIPS 2022 · 被引用 196 次
- Implicit Bias of AdamW: ℓ∞-Norm Constrained OptimizationShuo Xie, Zhiyuan LiICML 2024 · 被引用 46 次
相关 Paper
- The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural NetworksEitan Gronich, Gal VardiICML 2026
- Implicit Bias of Gradient Descent for Non-Homogeneous Deep NetworksYuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei 等ICML 2025
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural NetworksBohan Wang, Qi Meng, Wei Chen, Tie-Yan LiuICML 2021 · 被引用 45 次
- Implicit Bias of Adversarial Training for Deep Neural NetworksBochen Lv, Zhanxing ZhuICLR 2022 · 被引用 8 次
- The Implicit Bias of Adam on Separable DataChenyang Zhang, Difan Zou, Yuan CaoNeurIPS 2024 · 被引用 37 次
