Implicit Bias of AdamW: ℓ∞-Norm Constrained Optimization
Shuo Xie, Zhiyuan Li
Abstract
Adam with decoupled weight decay, also known as AdamW, is widely acclaimed for its superior performance in language modeling tasks, surpassing Adam with regularization in terms of generalization and optimization. However, this advantage is not theoretically well-understood. One challenge here is that though intuitively Adam with regularization optimizes the regularized loss, it is not clear if AdamW optimizes a specific objective. In this work, we make progress toward understanding the benefit of AdamW by showing that it implicitly performs constrained optimization. More concretely, we show in the full-batch setting, if AdamW converges with any non-increasing learning rate schedule whose partial sum diverges, it must converge to a KKT point of the original loss under the constraint that the norm of the parameter is bounded by the inverse of the weight decay factor. This result is built on the observation that Adam can be viewed as a smoothed version of SignGD, which is the normalized steepest descent with respect to norm, and a surprising connection between normalized steepest descent with weight decay and Frank-Wolfe.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers28
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed NoiseMaria-Eleni Sfyraki, Jun-Kun WangICML 2026 · 37 citations
- The Implicit Bias of Adam on Separable DataChenyang Zhang, Difan Zou, Yuan CaoNeurIPS 2024 · 37 citations
- Cautious Weight DecayLizhang Chen, Jonathan Li, Kaizhao Liang, Baiyu Su et al.ICLR 2026 · 14 citations
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi et al.ICLR 2026 · 14 citations
Builds on23
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
Related papers
- Never Saddle for Reparameterized Steepest Descent as Mirror FlowTom Jacobs, Chao Zhou, Rebekka BurkholzICLR 2026 · 3 citations
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 37 citations
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDTao Sun, Yuhao Huang, Li Shen, Kele Xu et al.CVPR 2025
- On the O(√d/K1/4) Convergence Rate of AdamW Measured by ℓ1 NormHuan Li, Yiming Dong, Zhouchen LinNeurIPS 2025 · 10 citations
- On the Overlooked Pitfalls of Weight Decay and How to Mitigate Them: A Gradient-Norm PerspectiveZeke Xie, Zhiqiang Xu, Jingzhao Zhang, Issei Sato et al.NeurIPS 2023 · 38 citations
