Stability and Generalization of Bilevel Programming in Hyperparameter Optimization
Fan Bao, Guoqiang Wu, Chongxuan Li, Jun Zhu, Bo Zhang
摘要
The (gradient-based) bilevel programming framework is widely used in hyperparameter optimization and has achieved excellent performance empirically. Previous theoretical work mainly focuses on its optimization properties, while leaving the analysis on generalization largely open. This paper attempts to address the issue by presenting an expectation bound w.r.t. the validation set based on uniform stability. Our results can explain some mysterious behaviours of the bilevel programming in practice, for instance, overfitting to the validation set. We also present an expectation bound for the classical cross-validation algorithm. Our results suggest that gradient-based algorithms can be better than cross-validation under certain conditions in a theoretical perspective. Furthermore, we prove that regularization terms in both the outer and inner levels can relieve the overfitting problem in gradient-based algorithms. In experiments on feature learning and data reweighting for noisy labels, we corroborate our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- A Fully First-Order Method for Stochastic Bilevel OptimizationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICML 2023 · 被引用 123 次
- Deep Safe Incomplete Multi-view Clustering: Theorem and AlgorithmHuayi Tang, Yong LiuICML 2022 · 被引用 118 次
- On Penalty Methods for Nonconvex Bilevel Optimization and First-Order Stochastic ApproximationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICLR 2024 · 被引用 61 次
- Stability and Generalization Analysis of Gradient Methods for Shallow Neural NetworksYunwen Lei, Rong Jin, Yiming YingNeurIPS 2022 · 被引用 30 次
- Subspace Learning for Effective Meta-LearningWeisen Jiang, James T. Kwok, Yu ZhangICML 2022 · 被引用 28 次
它引用的顶会 Paper5
- Bilevel Optimization: Convergence Analysis and Enhanced DesignKaiyi Ji, Junjie Yang, Yingbin LiangICML 2021 · 被引用 343 次
- Safe Deep Semi-Supervised Learning for Unseen-Class Unlabeled DataLan-Zhe Guo, Zhenyu Zhang, Yuan Jiang, Yufeng Li 等ICML 2020 · 被引用 243 次
- On the Iteration Complexity of Hypergradient ComputationRiccardo Grazzi, Luca Franceschi, Massimiliano Pontil, Saverio SalzoICML 2020 · 被引用 241 次
- Convergence of Meta-Learning with Task-Specific Adaptation over Partial ParametersKaiyi Ji, Jason D. Lee, Yingbin Liang, H. Vincent PoorNeurIPS 2020 · 被引用 97 次
- A Closer Look at the Training Strategy for Modern Meta-LearningJiaxin Chen, Xiao-Ming Wu, Yanke Li, Qimai Li 等NeurIPS 2020 · 被引用 48 次
相关 Paper
- Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter OptimizationRongzhen Wang, Chenyu Zheng, Guoqiang Wu, Xu Min 等NeurIPS 2024 · 被引用 3 次
- Convergence of Bayesian Bilevel OptimizationShi Fu, Fengxiang He, Xinmei Tian, Dacheng TaoICLR 2024 · 被引用 5 次
- Provably Data-driven Multiple Hyper-parameter Tuning with Structured Loss FunctionQUOC TUNG LE, Anh Nguyen, Viet Anh NguyenICML 2026 · 被引用 2 次
- Bilevel Optimization with Lower-Level Uniform Convexity: Theory and AlgorithmYuman Wu, Xiaochuan Gong, Jie Hao, Mingrui LiuICLR 2026 · 被引用 2 次
- A Generalized Weighted Optimization Method for Computational Learning and InversionKui Ren, Yunan Yang, Björn EngquistICLR 2022 · 被引用 4 次
