Gradient-based Hyperparameter Optimization Over Long Horizons
Paul Micaelli, Amos J. Storkey
Abstract
Gradient-based hyperparameter optimization has earned a widespread popularity in the context of few-shot meta-learning, but remains broadly impractical for tasks with long horizons (many gradient steps), due to memory scaling and gradient degradation issues. A common workaround is to learn hyperparameters online, but this introduces greediness which comes with a significant performance drop. We propose forward-mode differentiation with sharing (FDS), a simple and efficient algorithm which tackles memory scaling issues with forward-mode differentiation, and gradient degradation issues by sharing hyperparameters that are contiguous in time. We provide theoretical guarantees about the noise reduction properties of our algorithm, and demonstrate its efficiency empirically by differentiating through gradient steps of unrolled optimization. We consider large hyperparameter search ranges on CIFAR-10 where we significantly outperform greedy gradient-based alternatives, while achieving speedups compared to the state-of-the-art black-box methods. Code is available at: https://github.com/polo5/FDS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4415f169-c234-447d-8fb7-b26daf289b6eCited by top-tier papers10
- Bidirectional Learning for Offline Model-based Biological Sequence DesignCan Chen, Yingxue Zhang, Xue Liu, Mark CoatesICML 2023 · 30 citations
- Continuous-Time Meta-Learning with Forward Mode DifferentiationTristan Deleu, David Kanaa, Leo Feng, Giancarlo Kerg et al.ICLR 2022 · 22 citations
- Optimizing Canaries for Privacy Auditing with Metagradient DescentMatteo Boglioni, Terrance Liu, Andrew Ilyas, Steven WuICLR 2026 · 7 citations
- UQ-Guided Hyperparameter Optimization for Iterative LearnersJiesong Liu, Feng Zhang, Jiawei Guan, Xipeng ShenNeurIPS 2024 · 4 citations
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningMinyoung Kim, Timothy M. HospedalesAAAI 2025 · 3 citations
Builds on4
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Understanding and Robustifying Differentiable Architecture SearchArber Zela, Thomas Elsken, Tonmoy Saikia, Yassine Marrakchi et al.ICLR 2020 · 408 citations
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin et al.ICLR 2020 · 221 citations
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 58 citations
Related papers
- MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningBaoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li et al.AAAI 2024 · 91 citations
- Large-Scale Meta-Learning with Continual Trajectory ShiftingJaewoong Shin, Haebeom Lee, Boqing Gong, Sung Ju HwangICML 2021 · 18 citations
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang et al.ICLR 2022 · 6 citations
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter OptimizationOndrej Bohdal, Yongxin Yang, Timothy M. HospedalesNeurIPS 2021 · 29 citations
- Scalable Meta-Learning via Mixed-Mode DifferentiationIurii Kemaev, Dan A. Calian, Luisa M. Zintgraf, Gregory Farquhar et al.ICML 2025
