Lune

ICML2026Top-tier venue

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

Eitan Gronich, Gal Vardi

2026Year

Abstract

We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that momentum steepest descent algorithms like Muon (spectral norm), MomentumGD (ℓ2\ell_2 norm), and Signum (ℓ∞\ell_\infty norm) are approximate steepest descent trajectories under a decaying learning rate schedule, proving that these algorithms have a bias towards KKT points of the corresponding margin maximization problem. We extend the analysis to Adam (without the stability constant), which maximizes the ℓ∞\ell_\infty margin, and to Muon-Signum and Muon-Adam, which maximize a hybrid norm. Our experiments corroborate the theory and show that the identity of the margin maximized depends on the choice of optimizer. Overall, our results extend earlier lines of work on steepest descent in homogeneous models and momentum-based optimizers in linear models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 96e2b54e-313d-4dd7-a59b-dbc633f8b584

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines