Stochastic Gradient Descent for Gaussian Processes Done Right
Jihao Andreas Lin, Shreyas Padhy, Javier Antorán, Austin Tripp, Alexander Terenin, Csaba Szepesvári, José Miguel Hernández-Lobato, David Janz
Abstract
As is well known, both sampling from the posterior and computing the mean of the posterior in Gaussian process regression reduces to solving a large linear system of equations. We study the use of stochastic gradient descent for solving this linear system, and show that when done right -- by which we mean using specific insights from the optimisation and kernel communities -- stochastic gradient descent is highly effective. To that end, we introduce a particularly simple stochastic dual descent algorithm, explain its design in an intuitive manner and illustrate the design choices through a series of ablation studies. Further experiments demonstrate that our new method is highly competitive. In particular, our evaluations on the UCI regression tasks and on Bayesian optimisation set our approach apart from preconditioned conjugate gradients and variational Gaussian process approximations. Moreover, our method places Gaussian process regression on par with state-of-the-art graph neural networks for molecular binding affinity prediction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e47b1932-faf8-49e2-bb19-9b2b33adece3Cited by top-tier papers7
- Turbocharging Gaussian Process Inference with Approximate Sketch-and-ProjectPratik Rathore, Zachary Frangella, Sachin Garg, Shaghayegh Fazliani et al.NeurIPS 2025 · 8 citations
- Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian ProcessesJihao Andreas Lin, Shreyas Padhy, Bruno Mlodozeniec, Javier Antorán et al.NeurIPS 2024 · 6 citations
- Graph Random Features for Scalable Gaussian ProcessesMatthew Zhang, Jihao Andreas Lin, Krzysztof Choromanski, Adrian Weller et al.ICLR 2026 · 4 citations
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi et al.SIGCOMM 2026 · 1 citation
- The Price of Linear Time: Error Analysis of Structured Kernel InterpolationAlexander Moreno, Justin Xiao, Jonathan MeiICML 2025
Builds on5
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky et al.ICML 2020 · 186 citations
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 52 citations
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz et al.NeurIPS 2023 · 34 citations
- Tanimoto Random Features for Scalable Molecular Machine LearningAustin Tripp, Sergio Bacallado, Sukriti Singh, José Miguel Hernández-LobatoNeurIPS 2023 · 18 citations
- Sampling-based inference for large linear models, with application to linearised LaplaceJavier Antorán, Shreyas Padhy, Riccardo Barbano, Eric T. Nalisnick et al.ICLR 2023 · 1 citation
Related papers
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate GradientsArtem Artemev, David R. Burt, Mark van der WilkICML 2021 · 28 citations
- Scaling Gaussian Processes with Derivative Information Using Variational InferenceMisha Padidar, Xinran Zhu, Leo Huang, Jacob R. Gardner et al.NeurIPS 2021 · 28 citations
- Task-Agnostic Amortized Inference of Gaussian Process HyperparametersSulin Liu, Xingyuan Sun, Peter J. Ramadge, Ryan P. AdamsNeurIPS 2020 · 27 citations
- Adjoint-aided inference of Gaussian process driven differential equationsPaterne Gahungu, Christopher W. Lanyon, Mauricio A. Álvarez, Engineer Bainomugisha et al.NeurIPS 2022 · 6 citations
- Preconditioning for Scalable Gaussian Process Hyperparameter OptimizationJonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham et al.ICML 2022 · 36 citations
