Lune

ICLR2026Top-tier venue

Convex Dominance in Deep Learning I: A Scaling Law of Loss and Learning Rate

Zhiqi Bu, Shiyun Xu, Jialin Mao

2026Year
4Citations

Abstract

Deep learning has non-convex loss landscape and its optimization dynamics is hard to analyze or control. Nevertheless, the dynamics can be empirically convex-like across various tasks, models, optimizers, hyperparameters, etc. In this work, we examine the applicability of convexity and Lipschitz continuity in deep learning, in order to precisely control the loss dynamics via the learning rate schedules. We illustrate that deep learning quickly becomes weakly convex after a short period of training, and the loss is predicable by an upper bound on the last iterate, which further informs the scaling of optimal learning rate. Through the lens of convexity, we build scaling laws of learning rates and losses that extrapolate as much as 80×80\times across training horizons and 70×70\times across model sizes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bf8ccbdc-85c5-43b9-bedd-5a4901503cd8

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines