The Implicit Bias of Adam on Separable Data
Chenyang Zhang, Difan Zou, Yuan Cao
Abstract
Adam has become one of the most favored optimizers in deep learning problems. Despite its success in practice, numerous mysteries persist regarding its theoretical understanding. In this paper, we study the implicit bias of Adam in linear logistic regression. Specifically, we show that when the training data are linearly separable, Adam converges towards a linear classifier that achieves the maximum -margin. Notably, for a general class of diminishing learning rates, this convergence occurs within polynomial time. Our result shed light on the difference between Adam and (stochastic) gradient descent from a theoretical perspective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers23
- On Convergence of Adam for Stochastic Optimization under Relaxed AssumptionsYusu Hong, Junhong LinNeurIPS 2024 · 37 citations
- Implicit Optimization Bias of Next-token Prediction in Linear ModelsChristos ThrampoulidisNeurIPS 2024 · 19 citations
- How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced DataBhavya Vasudeva, Puneesh Deora, Yize Zhao, Vatsal Sharan et al.ICLR 2026 · 16 citations
- The Rich and the Simple: On the Implicit Bias of Adam and SGDBhavya Vasudeva, Jung Hoon Lee, Vatsal Sharan, Mahdi SoltanolkotabiNeurIPS 2025 · 14 citations
- Why is Your Language Model a Poor Implicit Reward Model?Noam Razin, Yong Lin, Jiarui Yao, Sanjeev AroraICLR 2026 · 8 citations
Builds on21
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
Related papers
- Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch RegimeBeomhan Baek, Minhak Song, Chulhee YunICLR 2026 · 2 citations
- The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural NetworksBohan Wang, Qi Meng, Wei Chen, Tie-Yan LiuICML 2021 · 45 citations
- Does Momentum Change the Implicit Regularization on Separable Data?Bohan Wang, Qi Meng, Huishuai Zhang, Ruoyu Sun et al.NeurIPS 2022 · 29 citations
- Implicit Bias of Spectal Descent and Muon on Multiclass Separable DataChen Fan, Mark Schmidt, Christos ThrampoulidisNeurIPS 2025
- Faster Margin Maximization Rates for Generic Optimization MethodsGuanghui Wang, Zihao Hu, Vidya Muthukumar, Jacob D. AbernethyNeurIPS 2023 · 3 citations
