Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes
Hao Chen, Lili Zheng, Raed Al Kontar, Garvesh Raskutti
摘要
Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full loss function and recovers model hyperparameters with rate O( 1 K ) up to a statistical error term depending on the minibatch size. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening up a new, previously unexplored, data size regime for GPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Few-Shot Bayesian Optimization with Deep Kernel SurrogatesMartin Wistuba, Josif GrabockaICLR 2021 · 被引用 87 次
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 被引用 38 次
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz 等NeurIPS 2023 · 被引用 34 次
- Using Random Effects to Account for High-Cardinality Categorical Features and Repeated Measures in Deep Neural NetworksGiora Simchoni, Saharon RossetNeurIPS 2021 · 被引用 27 次
- Existence and Estimation of Critical Batch Size for Training Generative Adversarial Networks with Two Time-Scale Update RuleNaoki Sato, Hideaki IidukaICML 2023 · 被引用 11 次
相关 Paper
- Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGDRémi Bardenet, Subhroshekhar Ghosh, Meixia LinNeurIPS 2021 · 被引用 13 次
- Sparse within Sparse Gaussian Processes using Neighbor InformationGia-Lac Tran, Dimitrios Milios, Pietro Michiardi, Maurizio FilipponeICML 2021 · 被引用 19 次
- Learning Curves for SGD on Structured FeaturesBlake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 29 次
- Computation-Aware Gaussian Processes: Model Selection And Linear-Time InferenceJonathan Wenger, Kaiwen Wu, Philipp Hennig, Jacob R. Gardner 等NeurIPS 2024 · 被引用 15 次
- Hybrid Stochastic-Deterministic Minibatch Proximal Gradient: Less-Than-Single-Pass Optimization with Nearly Optimal GeneralizationPan Zhou, Xiao-Tong YuanICML 2020 · 被引用 6 次
