Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes
Hao Chen, Lili Zheng, Raed Al Kontar, Garvesh Raskutti
Abstract
Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full loss function and recovers model hyperparameters with rate O( 1 K ) up to a statistical error term depending on the minibatch size. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening up a new, previously unexplored, data size regime for GPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b44c969-94e8-4278-95aa-a01201eccec1Cited by top-tier papers6
- Few-Shot Bayesian Optimization with Deep Kernel SurrogatesMartin Wistuba, Josif GrabockaICLR 2021 · 87 citations
- Global Convergence and Stability of Stochastic Gradient DescentVivak Patel, Shushu Zhang, Bowen TianNeurIPS 2022 · 38 citations
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz et al.NeurIPS 2023 · 34 citations
- Using Random Effects to Account for High-Cardinality Categorical Features and Repeated Measures in Deep Neural NetworksGiora Simchoni, Saharon RossetNeurIPS 2021 · 27 citations
- Existence and Estimation of Critical Batch Size for Training Generative Adversarial Networks with Two Time-Scale Update RuleNaoki Sato, Hideaki IidukaICML 2023 · 11 citations
Related papers
- Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGDRémi Bardenet, Subhroshekhar Ghosh, Meixia LinNeurIPS 2021 · 13 citations
- Sparse within Sparse Gaussian Processes using Neighbor InformationGia-Lac Tran, Dimitrios Milios, Pietro Michiardi, Maurizio FilipponeICML 2021 · 19 citations
- Learning Curves for SGD on Structured FeaturesBlake Bordelon, Cengiz PehlevanICLR 2022 · 29 citations
- Computation-Aware Gaussian Processes: Model Selection And Linear-Time InferenceJonathan Wenger, Kaiwen Wu, Philipp Hennig, Jacob R. Gardner et al.NeurIPS 2024 · 15 citations
- Hybrid Stochastic-Deterministic Minibatch Proximal Gradient: Less-Than-Single-Pass Optimization with Nearly Optimal GeneralizationPan Zhou, Xiao-Tong YuanICML 2020 · 6 citations
