Lune

ICLR2024Top-tier venue

Convergence of Bayesian Bilevel Optimization

Shi Fu, Fengxiang He, Xinmei Tian, Dacheng Tao

2024Year
5Citations
1Top-tier citations

Abstract

This paper presents the first theoretical guarantee for Bayesian bilevel optimization (BBO) that we term for the prevalent bilevel framework combining Bayesian optimization at the outer level to tune hyperparameters, and the inner-level stochastic gradient descent (SGD) for training the model. We prove sublinear regret bounds suggesting simultaneous convergence of the inner-level model parameters and outer-level hyperparameters to optimal configurations for generalization capability. A pivotal, technical novelty in the proofs is modeling the excess risk of the SGDtrained parameters as evaluation noise during Bayesian optimization. Our theory implies the inner unit horizon, defined as the number of SGD iterations, shapes the convergence behavior of BBO. This suggests practical guidance on configuring the inner unit horizon to enhance training efficiency and model performance. Hyperparameter optimization is crucial for leveraging deep learning's capabilities (Yang & Shami (2020); Elsken et al. (2019)). Techniques span Bayesian optimization (Wu et al. (2019); Victoria & Maragatham (2021)), decision theory (Bergstra & Bengio (2012)), multi-fidelity methods (Li et al. (2017)), and gradient-based approaches (Maclaurin et al. (2015)). We focus on BBO, exploring Bayesian optimization and bilevel frameworks' relevant aspects for hyperparameter tuning. Bayesian optimization. Bayesian optimization (BO) (Osborne & Osborne (2010); Kandasamy et al. (2020)) is a prevalent approach for hyperparameter tuning by efficiently exploring and exploiting hyperparameter spaces (Nguyen et al. (2017)). Gaussian processes (Bogunovic et al. (2018)) are commonly used as priors in BO to model uncertainty and estimate objective function distributions (Bro (2010); Wilson et al. (2014)). Among acquisition functions guiding queries in BO, the EI acquisition function (Jones & Welch (1998); Malkomes & Garnett (2018); Scarlett et al. (2017); Qin et al. (2017)) is one of the most widely utilized for balancing exploration-exploitation (Nguyen & Osborne (2020); Zhan & Xing (2020)). Other acquisitions like UCB (Valko et al. (2013)), knowledge

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3ff3dee5-df8b-45fa-93e6-cff3ba6e7222

Cited by top-tier papers1

Ask how each one uses it

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines