Random Function Descent
Felix Benning, Leif Döring
摘要
Classical worst-case optimization theory neither explains the success of optimization in machine learning, nor does it help with step size selection. In this paper we demonstrate the viability and advantages of replacing the classical 'convex function' framework with a 'random function' framework. With complexity , where is the number of steps and the number of dimensions, Bayesian optimization with gradients has not been viable in large dimension so far. By bridging the gap between Bayesian optimization (i.e. random function optimization theory) and classical optimization we establish viability. Specifically, we use a 'stochastic Taylor approximation' to rediscover gradient descent, which is scalable in high dimension due to complexity. This rediscovery yields a specific step size schedule we call Random Function Descent (RFD). The advantage of this random function framework is that RFD is scale invariant and that it provides a theoretical foundation for common step size heuristics such as gradient clipping and gradual learning rate warmup.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Learning-Rate-Free Learning by D-AdaptationAaron Defazio, Konstantin MishchenkoICML 2023 · 被引用 117 次
- Optimal Randomized First-Order Methods for Least-Squares ProblemsJonathan Lacotte, Mert PilanciICML 2020 · 被引用 30 次
- Scaling Gaussian Processes with Derivative Information Using Variational InferenceMisha Padidar, Xinran Zhu, Leo Huang, Jacob R. Gardner 等NeurIPS 2021 · 被引用 28 次
- High-Dimensional Gaussian Process Inference with DerivativesFilip de Roos, Alexandra Gessner, Philipp HennigICML 2021 · 被引用 24 次
相关 Paper
- On the Double Descent of Random Features Models Trained with SGDFanghui Liu, Johan A. K. Suykens, Volkan CevherNeurIPS 2022 · 被引用 11 次
- Re-Examining Linear Embeddings for High-Dimensional Bayesian OptimizationBenjamin Letham, Roberto Calandra, Akshara Rai, Eytan BakshyNeurIPS 2020 · 被引用 152 次
- To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-DimensionsNoah Marshall, Ke Liang Xiao, Atish Agarwala, Elliot PaquetteICLR 2025
- Trading Convergence Rate with Computational Budget in High Dimensional Bayesian OptimizationHung Tran-The, Sunil Gupta, Santu Rana, Svetha VenkateshAAAI 2020 · 被引用 14 次
- Vanilla Bayesian Optimization Performs Great in High DimensionsCarl Hvarfner, Erik Orm Hellsten, Luigi NardiICML 2024 · 被引用 88 次
