Amortized Proximal Optimization
Juhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. Grosse
Abstract
We propose a framework for online meta-optimization of parameters that govern optimization, called Amortized Proximal Optimization (APO). We first interpret various existing neural network optimizers as approximate stochastic proximal point methods which trade off the current-batch loss with proximity terms in both function space and weight space. The idea behind APO is to amortize the minimization of the proximal point objective by meta-learning the parameters of an update rule. We show how APO can be used to adapt a learning rate or a structured preconditioning matrix. Under appropriate assumptions, APO can recover existing optimizers such as natural gradient descent and KFAC. It enjoys low computational overhead and avoids expensive and numerically sensitive operations required by some second-order optimizers, such as matrix inverses. We empirically test APO for online adaptation of learning rates and structured preconditioning matrices for regression, image reconstruction, image classification, and natural language translation tasks. Empirically, the learning rate schedules found by APO generally outperform optimal fixed learning rates and are competitive with manually tuned decay schedules. Using APO to adapt a structured preconditioning matrix generally results in optimization performance competitive with second-order methods. Moreover, the absence of matrix inversion provides numerical stability, making it effective for low precision training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Training Data Attribution via Approximate UnrollingJuhan Bae, Wu Lin, Jonathan Lorraine, Roger B. GrosseNeurIPS 2024 · 41 citations
- Searching for Optimal Per-Coordinate Step-sizes with Multidimensional BacktrackingFrederik Kunstner, Victor Sanches Portella, Mark Schmidt, Nicholas J. A. HarveyNeurIPS 2023 · 14 citations
- Efficient Parametric Approximations of Neural Network Function Space DistanceNikita Dhawan, Sicong Huang, Juhan Bae, Roger Baker GrosseICML 2023 · 7 citations
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor GeneralizationDavide Buffelli, Jamie McGowan, Wangkun Xu, Alexandru Cioba et al.NeurIPS 2024 · 6 citations
- DISC: Dynamic Feature Selection for Cost-Sensitive Medical DiagnosisYusheng Li, Xincen Duan, Beili Wang, Wei Guo et al.AAAI 2026
Builds on5
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
- Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent GeneralizationNeha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer et al.ICML 2021 · 18 citations
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
Related papers
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 34 citations
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu et al.SC 2020 · 26 citations
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
- MADA: Meta-Adaptive Optimizers Through Hyper-Gradient DescentKaan Ozkara, Can Karakus, Parameswaran Raman, Mingyi Hong et al.ICML 2024 · 7 citations
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 31 citations
