Gradient-based Hyperparameter Optimization Over Long Horizons
Paul Micaelli, Amos J. Storkey
摘要
Gradient-based hyperparameter optimization has earned a widespread popularity in the context of few-shot meta-learning, but remains broadly impractical for tasks with long horizons (many gradient steps), due to memory scaling and gradient degradation issues. A common workaround is to learn hyperparameters online, but this introduces greediness which comes with a significant performance drop. We propose forward-mode differentiation with sharing (FDS), a simple and efficient algorithm which tackles memory scaling issues with forward-mode differentiation, and gradient degradation issues by sharing hyperparameters that are contiguous in time. We provide theoretical guarantees about the noise reduction properties of our algorithm, and demonstrate its efficiency empirically by differentiating through gradient steps of unrolled optimization. We consider large hyperparameter search ranges on CIFAR-10 where we significantly outperform greedy gradient-based alternatives, while achieving speedups compared to the state-of-the-art black-box methods. Code is available at: https://github.com/polo5/FDS
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Bidirectional Learning for Offline Model-based Biological Sequence DesignCan Chen, Yingxue Zhang, Xue Liu, Mark CoatesICML 2023 · 被引用 30 次
- Continuous-Time Meta-Learning with Forward Mode DifferentiationTristan Deleu, David Kanaa, Leo Feng, Giancarlo Kerg 等ICLR 2022 · 被引用 22 次
- Optimizing Canaries for Privacy Auditing with Metagradient DescentMatteo Boglioni, Terrance Liu, Andrew Ilyas, Steven WuICLR 2026 · 被引用 7 次
- UQ-Guided Hyperparameter Optimization for Iterative LearnersJiesong Liu, Feng Zhang, Jiawei Guan, Xipeng ShenNeurIPS 2024 · 被引用 4 次
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningMinyoung Kim, Timothy M. HospedalesAAAI 2025 · 被引用 3 次
它引用的顶会 Paper4
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Understanding and Robustifying Differentiable Architecture SearchArber Zela, Thomas Elsken, Tonmoy Saikia, Yassine Marrakchi 等ICLR 2020 · 被引用 408 次
- Meta-Learning with Warped Gradient DescentSebastian Flennerhag, Andrei A. Rusu, Razvan Pascanu, Francesco Visin 等ICLR 2020 · 被引用 221 次
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 被引用 58 次
相关 Paper
- MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningBaoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 等AAAI 2024 · 被引用 91 次
- Large-Scale Meta-Learning with Continual Trajectory ShiftingJaewoong Shin, Haebeom Lee, Boqing Gong, Sung Ju HwangICML 2021 · 被引用 18 次
- Online Hyperparameter Meta-Learning with Hypergradient DistillationHaebeom Lee, Hayeon Lee, Jaewoong Shin, Eunho Yang 等ICLR 2022 · 被引用 6 次
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter OptimizationOndrej Bohdal, Yongxin Yang, Timothy M. HospedalesNeurIPS 2021 · 被引用 29 次
- Scalable Meta-Learning via Mixed-Mode DifferentiationIurii Kemaev, Dan A. Calian, Luisa M. Zintgraf, Gregory Farquhar 等ICML 2025
