General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization
Kwangjun Ahn, Gagik Magakyan, Ashok Cutkosky
摘要
This work investigates the effectiveness of schedule-free methods, developed by A. Defazio et al. (NeurIPS 2024), in nonconvex optimization settings, inspired by their remarkable empirical success in training neural networks. Specifically, we show that schedule-free SGD achieves optimal iteration complexity for nonsmooth, nonconvex optimization problems. Our proof begins with the development of a general framework for online-to-nonconvex conversion, which converts a given online learning algorithm into an optimization algorithm for nonconvex losses. Our general framework not only recovers existing conversions but also leads to two novel conversion schemes. Notably, one of these new conversions corresponds directly to schedule-free SGD, allowing us to establish its optimality. Additionally, our analysis provides valuable insights into the parameter choices for schedule-free SGD, addressing a theoretical gap that the convex theory cannot explain. Introduction Training large-scale neural network models, such as large language models, requires a well-designed optimization strategy to ensure stable and fast convergence. For instance, training typically requires a carefully designed optimizer, such as the Adam optimizer [Kingma and Ba, 2014], along with meticulously tuned learning rate scheduling. Recently, Defazio et al. [2024] introduced the schedule-free method, which achieves impressive training performance without any need for learning rate scheduling. In brief, the schedule-free method is an add-on scheme that can be applied to any chosen base optimizer, converting it into a schedule-free variant. While this method has shown strong empirical performance in training large neural network models, its theoretical analysis has, to date, been limited to the convex setting [Defazio et al., 2024] . Our aim is to extend the theoretical understanding of schedule-free methods to nonconvex optimization. In this work, as an initial step, we focus on the version where the base optimizer is chosen as SGD, referred to as schedule-free SGD. For a given learning rate γ > 0 and interpolation weights c t , κ t ∈ [0, 1], ⋆ Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model TrainingMinhak Song, Beomhan Baek, Kwangjun Ahn, Chulhee YunNeurIPS 2025 · 被引用 9 次
- Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double OptimismFrancisco Patitucci, Ruichen Jiang, Aryan MokhtariICLR 2026 · 被引用 3 次
- FOAM: Frequency and Operator-Error Based Adaptive Damping Method for Reducing Staleness-Oriented Error for ShampooKyunghun Nam, Sumyeong AhnICML 2026
- Derandomized Online-to-Non-convex Conversion for Stochastic Weakly Convex OptimizationFanfan Ji, Xiaotong YuanICLR 2026
它引用的顶会 Paper11
- The Road Less ScheduledAaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko 等NeurIPS 2024 · 被引用 208 次
- Gradient-Free Methods for Deterministic and Stochastic Nonsmooth Nonconvex OptimizationTianyi Lin, Zeyu Zheng, Michael I. JordanNeurIPS 2022 · 被引用 102 次
- Complexity of Finding Stationary Points of Nonconvex Nonsmooth FunctionsJingzhao Zhang, Hongzhou Lin, Stefanie Jegelka, Suvrit Sra 等ICML 2020 · 被引用 98 次
- A gradient sampling method with complexity guarantees for Lipschitz functions in high and low dimensionsDamek Davis, Dmitriy Drusvyatskiy, Yin Tat Lee, Swati Padmanabhan 等NeurIPS 2022 · 被引用 77 次
- Optimal Stochastic Non-smooth Non-convex Optimization through Online-to-Non-convex ConversionAshok Cutkosky, Harsh Mehta, Francesco OrabonaICML 2023 · 被引用 54 次
相关 Paper
- Oracle Complexity in Nonsmooth Nonconvex OptimizationGuy Kornowski, Ohad ShamirNeurIPS 2021 · 被引用 74 次
- Random Scaling and Momentum for Non-smooth Non-convex OptimizationQinzi Zhang, Ashok CutkoskyICML 2024 · 被引用 10 次
- A Second look at Exponential and Cosine Step Sizes: Simplicity, Adaptivity, and PerformanceXiaoyu Li, Zhenxun Zhuang, Francesco OrabonaICML 2021 · 被引用 29 次
- The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model TrainingFabian Schaipp, Alexander Hägele, Adrien B. Taylor, Umut Simsekli 等ICML 2025
- Accelerating neural network training: An analysis of the AlgoPerf competitionPriya Kasimbeg, Frank Schneider, Runa Eschenhagen, Juhan Bae 等ICLR 2025
