Lune

ICLR2025

Adam Exploits ℓ∞-geometry of Loss Landscape via Coordinate-wise Adaptivity

Shuo Xie, Mohamad Amin Mohamadi, Zhiyuan Li

2025Year

Abstract

Adam outperforms SGD when training language models. Yet this advantage is not well-understood theoretically -previous convergence analysis for Adam and SGD mainly focuses on the number of steps T and is already minimax-optimal in non-convex cases, which are both O(T -1/4 ). In this work, we argue that the exploitation of nice ℓ∞-geometry is the key advantage of Adam over SGD. More specifically, we give a new convergence analysis for Adam under novel assumptions that loss is smooth under ℓ∞-geometry rather than the more common ℓ2-geometry, which yields a much better empirical smoothness constant for GPT-2 and ResNet models. Our experiments confirm that Adam performs much worse when the favorable ℓ∞-geometry is changed while SGD provably remains unaffected. We also extend the convergence analysis to blockwise Adam under novel blockwise smoothness assumptions.

L(x), we consider optimization over L with access only to independent stochastic functions L t (x) T t=1 such that EL t (x) = L(x) for any input x ∈ R d .

For an invertible function T : R d → R d , T is a rotating transformation if there exists an orthogonal matrix T ∈ R d×d such that T (x) = T x. T is a permutating transformation if there exists a permutation π : [d] → [d] such that T (x) = [x π(1) , . . . , x π(d) ] ⊤ . A permutating transformation is always a rotating transformation. We will use R to denote a rotating transformation. Definition 2.1. For initialization x 0 and stochastic losses L t T t=1 , we can get x t when running algorithm A on (x 0 , L t T t=1 ). For a transformation T , we can also get xt when running A with the same hyperparameters on ( x0 , Lt T t=1 ) with x0 = T -1 (x 0 ) and Lt = L t • T . An algorithm A is equivariant w.r.t. T if it always holds that xt = T -1 (x t ) for any hyperparameters, initialization and stochastic losses. An algorithm A is rotation-equivariant if it is equivariant w.r.t. any rotating transformation R. And A is permutation-equivariant if it is equivariant w.r.t. any permutating transformation.

The following Theorem 2.2 shows the difference between Adam and AdaSGD, whose proof is in Appendix A. We provide a visual example in Figure 2. It also shows the similarity between SignGD and Adam.