Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit Differentiation
Ross M. Clarke, Elre Talea Oldewage, José Miguel Hernández-Lobato
摘要
Machine learning training methods depend plentifully and intricately on hyperparameters, motivating automated strategies for their optimisation. Many existing algorithms restart training for each new hyperparameter choice, at considerable computational cost. Some hypergradient-based one-pass methods exist, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or take several times longer to train than their base models. We extend these existing methods to develop an approximate hypergradient-based hyperparameter optimiser which is applicable to any continuous hyperparameter appearing in a differentiable model weight update, yet requires only one training episode, with no restarts. We also provide a motivating argument for convergence to the true hypergradient, and perform tractable gradient-based optimisation of independent learning rates for each model parameter. Our method performs competitively from varied random hyperparameter initialisations on several UCI datasets and Fashion-MNIST (using a one-layer MLP), Penn Treebank (using an LSTM) and CIFAR-10 (using a ResNet-18), in time only 2-3x greater than vanilla training. INTRODUCTION Many machine learning methods are governed by hyperparameters: quantities other than model parameters or weights which nonetheless influence training (e.g. optimiser settings, dropout probabilities and dataset configurations). As suitable hyperparameter selection is crucial to system performance (e.g. Kohavi & John (1995) ), it is a pillar of efforts to automate machine learning (Hutter et al., 2018 , Chapter 1), spawning several hyperparameter optimisation (HPO) algorithms (e.g. Bergstra & Bengio (2012); Snoek et al. (2012; 2015); Falkner et al. (2018)). However, HPO is computationally intensive and random search is an unexpectedly strong (but beatable; Turner et al. (2021)) baseline; beyond random or grid searches, HPO is relatively underused in research (Bouthillier & Varoquaux, 2020). Recently, Lorraine et al. (2020) used gradient-based updates to adjust hyperparameters during training, displaying impressive optimisation performance and scalability to high-dimensional hyperparameters. Despite their computational efficiency (since updates occur before final training performance is known), Lorraine et al.'s algorithm only applies to hyperparameters on which the loss function depends explicitly (such as 2 regularisation), notably excluding optimiser hyperparameters. Our work extends Lorraine et al.'s algorithm to support arbitrary continuous inputs to a differentiable weight update formula, including learning rates and momentum factors. We demonstrate our algorithm handles a range of hyperparameter initialisations and datasets, improving test loss after a single training episode ('one pass'). Relaxing differentiation-through-optimisation (Domke, 2012) and hypergradient descent's (Baydin et al., 2018) exactness allows us to improve computational and memory efficiency. Our scalable one-pass method improves performance from arbitrary hyperparameter initialisations, and could be augmented with a further search over those initialisations if desired. WEIGHT-UPDATE HYPERPARAMETER TUNING In this section, we develop our method. Expanded derivations and a summary of differences from Lorraine et al. ( 2020 ) are given in Appendix C.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Making Scalable Meta Learning PracticalSang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger 等NeurIPS 2023 · 被引用 28 次
- Module-Aware Optimization for Auxiliary LearningHong Chen, Xin Wang, Yue Liu, Yuwei Zhou 等NeurIPS 2022 · 被引用 11 次
- Meta-learning Adaptive Deep Kernel Gaussian Processes for Molecular Property PredictionWenlin Chen, Austin Tripp, José Miguel Hernández-LobatoICLR 2023 · 被引用 8 次
- Studying K-FAC Heuristics by Viewing Adam through a Second-Order LensRoss M. Clarke, José Miguel Hernández-LobatoICML 2024 · 被引用 2 次
- Bi-level Physics-Informed Neural Networks for PDE Constrained Optimization using Broyden's HypergradientsZhongkai Hao, Chengyang Ying, Hang Su, Jun Zhu 等ICLR 2023 · 被引用 2 次
它引用的顶会 Paper1
相关 Paper
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 被引用 66 次
- Improved Penalty Method via Doubly Stochastic Gradients for Bilevel Hyperparameter OptimizationWanli Shi, Bin GuAAAI 2021 · 被引用 6 次
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parametersArsalan Sharifnassab, Saber Salehkaleybar, Richard S. SuttonICML 2025
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li 等NeurIPS 2021 · 被引用 111 次
- Frugal Optimization for Cost-related HyperparametersQingyun Wu, Chi Wang, Silu HuangAAAI 2021 · 被引用 51 次
