Scalable One-Pass Optimisation of High-Dimensional Weight-Update Hyperparameters by Implicit Differentiation
Ross M. Clarke, Elre Talea Oldewage, José Miguel Hernández-Lobato
Abstract
Machine learning training methods depend plentifully and intricately on hyperparameters, motivating automated strategies for their optimisation. Many existing algorithms restart training for each new hyperparameter choice, at considerable computational cost. Some hypergradient-based one-pass methods exist, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or take several times longer to train than their base models. We extend these existing methods to develop an approximate hypergradient-based hyperparameter optimiser which is applicable to any continuous hyperparameter appearing in a differentiable model weight update, yet requires only one training episode, with no restarts. We also provide a motivating argument for convergence to the true hypergradient, and perform tractable gradient-based optimisation of independent learning rates for each model parameter. Our method performs competitively from varied random hyperparameter initialisations on several UCI datasets and Fashion-MNIST (using a one-layer MLP), Penn Treebank (using an LSTM) and CIFAR-10 (using a ResNet-18), in time only 2-3x greater than vanilla training. INTRODUCTION Many machine learning methods are governed by hyperparameters: quantities other than model parameters or weights which nonetheless influence training (e.g. optimiser settings, dropout probabilities and dataset configurations). As suitable hyperparameter selection is crucial to system performance (e.g. Kohavi & John (1995) ), it is a pillar of efforts to automate machine learning (Hutter et al., 2018 , Chapter 1), spawning several hyperparameter optimisation (HPO) algorithms (e.g. Bergstra & Bengio (2012); Snoek et al. (2012; 2015); Falkner et al. (2018)). However, HPO is computationally intensive and random search is an unexpectedly strong (but beatable; Turner et al. (2021)) baseline; beyond random or grid searches, HPO is relatively underused in research (Bouthillier & Varoquaux, 2020). Recently, Lorraine et al. (2020) used gradient-based updates to adjust hyperparameters during training, displaying impressive optimisation performance and scalability to high-dimensional hyperparameters. Despite their computational efficiency (since updates occur before final training performance is known), Lorraine et al.'s algorithm only applies to hyperparameters on which the loss function depends explicitly (such as 2 regularisation), notably excluding optimiser hyperparameters. Our work extends Lorraine et al.'s algorithm to support arbitrary continuous inputs to a differentiable weight update formula, including learning rates and momentum factors. We demonstrate our algorithm handles a range of hyperparameter initialisations and datasets, improving test loss after a single training episode ('one pass'). Relaxing differentiation-through-optimisation (Domke, 2012) and hypergradient descent's (Baydin et al., 2018) exactness allows us to improve computational and memory efficiency. Our scalable one-pass method improves performance from arbitrary hyperparameter initialisations, and could be augmented with a further search over those initialisations if desired. WEIGHT-UPDATE HYPERPARAMETER TUNING In this section, we develop our method. Expanded derivations and a summary of differences from Lorraine et al. ( 2020 ) are given in Appendix C.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9804c2b4-7924-48a7-816a-6659669b0fe8Cited by top-tier papers5
- Making Scalable Meta Learning PracticalSang Keun Choe, Sanket Vaibhav Mehta, Hwijeen Ahn, Willie Neiswanger et al.NeurIPS 2023 · 28 citations
- Module-Aware Optimization for Auxiliary LearningHong Chen, Xin Wang, Yue Liu, Yuwei Zhou et al.NeurIPS 2022 · 11 citations
- Meta-learning Adaptive Deep Kernel Gaussian Processes for Molecular Property PredictionWenlin Chen, Austin Tripp, José Miguel Hernández-LobatoICLR 2023 · 8 citations
- Studying K-FAC Heuristics by Viewing Adam through a Second-Order LensRoss M. Clarke, José Miguel Hernández-LobatoICML 2024 · 2 citations
- Bi-level Physics-Informed Neural Networks for PDE Constrained Optimization using Broyden's HypergradientsZhongkai Hao, Chengyang Ying, Hang Su, Jun Zhu et al.ICLR 2023 · 2 citations
Builds on1
Related papers
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 66 citations
- Improved Penalty Method via Doubly Stochastic Gradients for Bilevel Hyperparameter OptimizationWanli Shi, Bin GuAAAI 2021 · 6 citations
- MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parametersArsalan Sharifnassab, Saber Salehkaleybar, Richard S. SuttonICML 2025
- Federated Hyperparameter Tuning: Challenges, Baselines, and Connections to Weight-SharingMikhail Khodak, Renbo Tu, Tian Li, Liam Li et al.NeurIPS 2021 · 111 citations
- Frugal Optimization for Cost-related HyperparametersQingyun Wu, Chi Wang, Silu HuangAAAI 2021 · 51 citations
