The Statistical Complexity of Early-Stopped Mirror Descent
Tomas Vaskevicius, Varun Kanade, Patrick Rebeschini
Abstract
Recently there has been a surge of interest in understanding implicit regularization properties of iterative gradient-based optimization algorithms. In this paper, we study the statistical guarantees on the excess risk achieved by early-stopped unconstrained mirror descent algorithms applied to the unregularized empirical risk. We consider the set-up of learning linear models and kernel methods for strongly convex and Lipschitz loss functions while imposing only boundedness conditions on the unknown data-generating mechanism. By completing an inequality that characterizes convexity for the squared loss, we identify an intrinsic link between offset Rademacher complexities and potential-based convergence analysis of mirror descent methods. Our observation immediately yields excess risk guarantees for the path traced by the iterates of mirror descent in terms of offset complexities of certain function classes depending only on the choice of the mirror map, initialization point, step size and the number of iterations. We apply our theory to recover, in a clean and elegant manner via rather short proofs, some of the recent results in the implicit regularization literature while also showing how to improve upon them in some settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 467c3fa6-70e3-4fb5-80d2-765b35bdc4a7Cited by top-tier papers9
- Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of StochasticityScott Pesme, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 135 citations
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 42 citations
- Sobolev Acceleration and Statistical Optimality for Learning Elliptic Equations via Gradient DescentYiping Lu, José H. Blanchet, Lexing YingNeurIPS 2022 · 15 citations
- Implicit Regularization in Matrix Sensing via Mirror DescentFan Wu, Patrick RebeschiniNeurIPS 2021 · 13 citations
- Thinking Outside the Ball: Optimal Learning with Gradient Descent for Generalized Linear Stochastic Convex OptimizationIdan Amir, Roi Livni, Nati SrebroNeurIPS 2022 · 7 citations
Builds on1
Related papers
- Localization, Convexity, and Star AggregationSuhas VijaykumarNeurIPS 2021 · 10 citations
- Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic RegressionJingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin YuICML 2025
- Never Go Full Batch (in Stochastic Convex Optimization)Idan Amir, Yair Carmon, Tomer Koren, Roi LivniNeurIPS 2021 · 17 citations
- Characterization of Excess Risk for Locally Strongly Convex Population RiskMingyang Yi, Ruoyu Wang, Zhi-Ming MaNeurIPS 2022 · 4 citations
- Mirror Descent Maximizes Generalized Margin and Can Be Implemented EfficientlyHaoyuan Sun, Kwangjun Ahn, Christos Thrampoulidis, Navid AzizanNeurIPS 2022 · 33 citations
