Two Losses Are Better Than One: Faster Optimization Using a Cheaper Proxy
Blake E. Woodworth, Konstantin Mishchenko, Francis R. Bach
摘要
We present an algorithm for minimizing an objective with hard-to-compute gradients by using a related, easier-to-access function as a proxy. Our algorithm is based on approximate proximal point iterations on the proxy combined with relatively few stochastic gradients from the objective. When the difference between the objective and the proxy is -smooth, our algorithm guarantees convergence at a rate matching stochastic gradient descent on a -smooth objective, which can lead to substantially better sample efficiency. Our algorithm has many potential applications in machine learning, and provides a principled means of leveraging synthetic data, physics simulators, mixed public and private data, and more.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Non-convex Stochastic Composite Optimization with Polyak MomentumYuan Gao, Anton Rodomanov, Sebastian U. StichICML 2024 · 被引用 13 次
- Accelerated Methods with Compressed Communications for Distributed Optimization Problems Under Data SimilarityDmitry Bylinkin, Aleksandr BeznosikovAAAI 2025 · 被引用 3 次
它引用的顶会 Paper4
- Rethinking ImageNet Pre-TrainingKaiming He, Ross B. Girshick, Piotr DollárICCV 2019 · 被引用 1,188 次
- Mind the Gap: Assessing Temporal Generalization in Neural Language ModelsAngeliki Lazaridou, Adhiguna Kuncoro, Elena Gribovskaya, Devang Agrawal 等NeurIPS 2021 · 被引用 315 次
- Statistically Preconditioned Accelerated Gradient Method for Distributed OptimizationHadrien Hendrikx, Lin Xiao, Sébastien Bubeck, Francis R. Bach 等ICML 2020 · 被引用 66 次
- Fast Stochastic Bregman Gradient Methods: Sharp Analysis and Variance ReductionRadu-Alexandru Dragomir, Mathieu Even, Hadrien HendrikxICML 2021 · 被引用 40 次
相关 Paper
- Target-based Surrogates for Stochastic OptimizationJonathan Wilder Lavington, Sharan Vaswani, Reza Babanezhad Harikandeh, Mark Schmidt 等ICML 2023 · 被引用 6 次
- Super-efficiency of automatic differentiation for functions defined as a minimumPierre Ablin, Gabriel Peyré, Thomas MoreauICML 2020 · 被引用 42 次
- Stochastic Approximate Gradient Descent via the Langevin AlgorithmYixuan Qiu, Xiao WangAAAI 2020 · 被引用 5 次
- Lower Complexity Bounds for Finite-Sum Convex-Concave Minimax Optimization ProblemsGuangzeng Xie, Luo Luo, Yijiang Lian, Zhihua ZhangICML 2020 · 被引用 21 次
- Descent with Misaligned Gradients and Applications to Hidden ConvexityAditya Bhaskara, Ashok Cutkosky, Ravi Kumar, Manish PurohitICLR 2025
