Amortized Proximal Optimization
Juhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. Grosse
摘要
We propose a framework for online meta-optimization of parameters that govern optimization, called Amortized Proximal Optimization (APO). We first interpret various existing neural network optimizers as approximate stochastic proximal point methods which trade off the current-batch loss with proximity terms in both function space and weight space. The idea behind APO is to amortize the minimization of the proximal point objective by meta-learning the parameters of an update rule. We show how APO can be used to adapt a learning rate or a structured preconditioning matrix. Under appropriate assumptions, APO can recover existing optimizers such as natural gradient descent and KFAC. It enjoys low computational overhead and avoids expensive and numerically sensitive operations required by some second-order optimizers, such as matrix inverses. We empirically test APO for online adaptation of learning rates and structured preconditioning matrices for regression, image reconstruction, image classification, and natural language translation tasks. Empirically, the learning rate schedules found by APO generally outperform optimal fixed learning rates and are competitive with manually tuned decay schedules. Using APO to adapt a structured preconditioning matrix generally results in optimization performance competitive with second-order methods. Moreover, the absence of matrix inversion provides numerical stability, making it effective for low precision training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 被引用 41 次
- Searching for Optimal Per-Coordinate Step-sizes with Multidimensional BacktrackingFrederik Kunstner, Victor Sanches Portella, Mark Schmidt, Nicholas J. A. HarveyNeurIPS 2023 · 被引用 14 次
- Efficient Parametric Approximations of Neural Network Function Space DistanceNikita Dhawan, Sicong Huang, Juhan Bae, Roger Baker GrosseICML 2023 · 被引用 7 次
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor GeneralizationDavide Buffelli, Jamie McGowan, Wangkun Xu, Alexandru Cioba 等NeurIPS 2024 · 被引用 6 次
- DISC: Dynamic Feature Selection for Cost-Sensitive Medical DiagnosisYusheng Li, Xincen Duan, Beili Wang, Wei Guo 等AAAI 2026
它引用的顶会 Paper5
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin 等ICLR 2020 · 被引用 221 次
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu 等ACL 2020 · 被引用 148 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent GeneralizationNeha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer 等ICML 2021 · 被引用 18 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
相关 Paper
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 被引用 34 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov 等ICML 2024 · 被引用 7 次
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong 等ICML 2024 · 被引用 7 次
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 被引用 31 次
