Lune

ICML2026Top-tier venue

Gradient Descent as a Perceptron Algorithm: Understanding Dynamics and Implicit Acceleration

Alexander Tyurin

2026Year

Abstract

Even for the gradient descent (GD) method applied to neural network training, understanding its optimization dynamics, including convergence rate, iterate trajectories, function value oscillations, and especially its implicit acceleration, remains a challenging problem. We analyze nonlinear models with the logistic loss and show that the steps of GD reduce to those of generalized perceptron algorithms (Rosenblatt, 1958), providing a new perspective on the dynamics. This reduction yields significantly simpler algorithmic steps, which we analyze using classical linear algebra tools. Using these tools, we demonstrate on a minimalistic example that the nonlinearity in a two-layer model can provably yield a faster iteration complexity O~(d)\tilde{\mathcal{O}}(\sqrt{d}) compared to Ω(d)\Omega(d) achieved by linear models, where dd is the number of features. This helps explain the optimization dynamics and the implicit acceleration phenomenon observed in neural networks. The theoretical results are supported by extensive numerical experiments. We believe that this alternative view will further advance research on the optimization of neural networks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cc5d59f4-2ccd-49d9-85ed-9ecb97e72461

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines