Stable Nonconvex-Nonconcave Training via Linear Interpolation
Thomas Pethick, Wanyun Xie, Volkan Cevher
摘要
This paper presents a theoretical analysis of linear interpolation as a principled method for stabilizing (large-scale) neural network training. We argue that instabilities in the optimization process are often caused by the nonmonotonicity of the loss landscape and show how linear interpolation can help by leveraging the theory of nonexpansive operators. We construct a new optimization scheme called relaxed approximate proximal point (RAPP), which is the first explicit method without anchoring to achieve last iterate convergence rates for -comonotone problems while only requiring . The construction extends to constrained and regularized settings. By replacing the inner optimizer in RAPP we rediscover the family of Lookahead algorithms for which we establish convergence in cohypomonotone problems even when the base optimizer is taken to be gradient descent ascent. The range of cohypomonotone problems in which Lookahead converges is further expanded by exploiting that Lookahead inherits the properties of the base optimizer. We corroborate the results with experiments on generative adversarial networks which demonstrates the benefits of the linear interpolation present in both RAPP and Lookahead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Accelerated Algorithms for Constrained Nonconvex-Nonconcave Min-Max Optimization and Comonotone InclusionYang Cai, Argyris Oikonomou, Weiqiang ZhengICML 2024 · 被引用 26 次
- Revisiting Inexact Fixed-Point Iterations for Min-Max Problems: Stochasticity and Structured NonconvexityAhmet Alacaoglu, Donghwan Kim, Stephen J. WrightICML 2024 · 被引用 6 次
- Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable ImprovementsJiseok Chae, Chulhee Yun, Donghwan KimNeurIPS 2024 · 被引用 2 次
- Last-Iterate Convergence of Regularized Gradient Methods for Stochastic Monotone Variational InequalitiesShinji Ito, Taira Tsuchiya, Kaito Ariu, Kenshi AbeICML 2026
- Efficient Interpolation between Extragradient and Proximal Methods for Weak MVIsThomas Pethick, Ioannis Mavrothalassitis, Volkan CevherICLR 2025
它引用的顶会 Paper8
- Accelerated Algorithms for Smooth Convex-Concave Minimax Problems with O(1/k^2) Rate on Squared Gradient NormTaeho Yoon, Ernest K. RyuICML 2021 · 被引用 138 次
- Fast Extra Gradient Methods for Smooth Structured Nonconvex-Nonconcave Minimax ProblemsSucheol Lee, Donghwan KimNeurIPS 2021 · 被引用 125 次
- The Limits of Min-Max Optimization Algorithms: Convergence to Spurious Non-Critical SetsYa-Ping Hsieh, Panayotis Mertikopoulos, Volkan CevherICML 2021 · 被引用 96 次
- Escaping limit cycles: Global convergence for constrained nonconvex-nonconcave minimax problemsThomas Pethick, Puya Latafat, Panos Patrinos, Olivier Fercoq 等ICLR 2022 · 被引用 60 次
- Towards Understanding Why Lookahead Generalizes Better Than SGD and BeyondPan Zhou, Hanshu Yan, Xiaotong Yuan, Jiashi Feng 等NeurIPS 2021 · 被引用 37 次
相关 Paper
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt 等ICML 2026
- On Monotonic Linear Interpolation of Neural Network ParametersJames Lucas, Juhan Bae, Michael R. Zhang, Stanislav Fort 等ICML 2021 · 被引用 11 次
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 被引用 15 次
- Sharpness-Aware Minimization Can Hallucinate MinimizersChanwoong Park, Uijeong Jang, Ernest Ryu, Insoon YangICML 2026
- Stacey: Promoting Stochastic Steepest Descent via Accelerated ℓp-Smooth Nonconvex OptimizationXinyu Luo, Site Bai, Bolian Li, Petros Drineas 等ICML 2025
