Towards Understanding Why Lookahead Generalizes Better Than SGD and Beyond
Pan Zhou, Hanshu Yan, Xiaotong Yuan, Jiashi Feng, Shuicheng Yan
Abstract
To train networks, lookahead algorithm [1] updates its fast weights k times via an inner-loop optimizer before updating its slow weights once by using the latest fast weights. Any optimizer, e.g. SGD, can serve as the inner-loop optimizer, and the derived lookahead generally enjoys remarkable test performance improvement over the vanilla optimizer. But theoretical understandings on the test performance improvement of lookahead remain absent yet. To solve this issue, we theoretically justify the advantages of lookahead in terms of the excess risk error which measures the test performance. Specifically, we prove that lookahead using SGD as its inner-loop optimizer can better balance the optimization error and generalization error to achieve smaller excess risk error than vanilla SGD on (strongly) convex problems and nonconvex problems with Polyak-Łojasiewicz condition which has been observed/proved in neural networks. Moreover, we show the stagewise optimization strategy [2] which decays learning rate several times during training can also benefit lookahead in improving its optimization and generalization errors on strongly convex problems. Finally, we propose a stagewise locally-regularized lookahead (SLRLA) algorithm which sums up the vanilla objective and a local regularizer to minimize at each stage and provably enjoys optimization and generalization improvement over the conventional (stagewise) lookahead. Experimental results on CIFAR10/100 and ImageNet testify its advantages. Codes is available at https://github.com/sail-sg/SLRLA-optimizer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Efficient Sharpness-aware Minimization for Improved Training of Neural NetworksJiawei Du, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou et al.ICLR 2022 · 168 citations
- Understanding How Consistency Works in Federated Learning via Stage-wise Relaxed InitializationYan Sun, Li Shen, Dacheng TaoNeurIPS 2023 · 26 citations
- Stable Nonconvex-Nonconcave Training via Linear InterpolationThomas Pethick, Wanyun Xie, Volkan CevherNeurIPS 2023 · 8 citations
- Understanding Outer Optimizers in Local SGD: Learning Rates, Momentum, and AccelerationAhmed Khaled, Satyen Kale, Arthur Douillard, Chi Jin et al.NeurIPS 2025 · 7 citations
- A-FedPD: Aligning Dual-Drift is All Federated Primal-Dual Learning NeedsYan Sun, Li Shen, Dacheng TaoNeurIPS 2024 · 6 citations
Builds on8
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- Prototypical Contrastive Learning of Unsupervised RepresentationsJunnan Li, Pan Zhou, Caiming Xiong, Steven C. H. HoiICLR 2021 · 484 citations
- Towards Theoretically Understanding Why Sgd Generalizes Better Than Adam in Deep LearningPan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong et al.NeurIPS 2020 · 309 citations
- Theory-Inspired Path-Regularized Differential Network Architecture SearchPan Zhou, Caiming Xiong, Richard Socher, Steven Chu-Hong HoiNeurIPS 2020 · 64 citations
Related papers
- Bidirectional Looking with A Novel Double Exponential Moving Average to Adaptive and Non-adaptive Momentum OptimizersYineng Chen, Zuchao Li, Lefei Zhang, Bo Du et al.ICML 2023 · 8 citations
- Lookaround Optimizer: k steps around, 1 step averageJiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu et al.NeurIPS 2023 · 12 citations
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDTao Sun, Yuhao Huang, Li Shen, Kele Xu et al.CVPR 2025
- STL-SGD: Speeding Up Local SGD with Stagewise Communication PeriodShuheng Shen, Yifei Cheng, Jingchang Liu, Linli XuAAAI 2021 · 12 citations
- Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial TimeSungyoon Kim, Mert PilanciICML 2024 · 10 citations
