Stochastic Gradient Descent for Gaussian Processes Done Right
Jihao Andreas Lin, Shreyas Padhy, Javier Antorán, Austin Tripp, Alexander Terenin, Csaba Szepesvári, José Miguel Hernández-Lobato, David Janz
摘要
As is well known, both sampling from the posterior and computing the mean of the posterior in Gaussian process regression reduces to solving a large linear system of equations. We study the use of stochastic gradient descent for solving this linear system, and show that when done right -- by which we mean using specific insights from the optimisation and kernel communities -- stochastic gradient descent is highly effective. To that end, we introduce a particularly simple stochastic dual descent algorithm, explain its design in an intuitive manner and illustrate the design choices through a series of ablation studies. Further experiments demonstrate that our new method is highly competitive. In particular, our evaluations on the UCI regression tasks and on Bayesian optimisation set our approach apart from preconditioned conjugate gradients and variational Gaussian process approximations. Moreover, our method places Gaussian process regression on par with state-of-the-art graph neural networks for molecular binding affinity prediction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Turbocharging Gaussian Process Inference with Approximate Sketch-and-ProjectPratik Rathore, Zachary Frangella, Sachin Garg, Shaghayegh Fazliani 等NeurIPS 2025 · 被引用 8 次
- Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian ProcessesJihao Andreas Lin, Shreyas Padhy, Bruno Mlodozeniec, Javier Antorán 等NeurIPS 2024 · 被引用 6 次
- Graph Random Features for Scalable Gaussian ProcessesMatthew Zhang, Jihao Andreas Lin, Krzysztof Choromanski, Adrian Weller 等ICLR 2026 · 被引用 4 次
- ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi 等SIGCOMM 2026 · 被引用 1 次
- The Price of Linear Time: Error Analysis of Structured Kernel InterpolationAlexander Moreno, Justin Xiao, Jonathan MeiICML 2025
它引用的顶会 Paper5
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky 等ICML 2020 · 被引用 186 次
- Last iterate convergence of SGD for Least-Squares in the Interpolation regimeAditya Vardhan Varre, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2021 · 被引用 52 次
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz 等NeurIPS 2023 · 被引用 34 次
- Tanimoto Random Features for Scalable Molecular Machine LearningAustin Tripp, Sergio Bacallado, Sukriti Singh, José Miguel Hernández-LobatoNeurIPS 2023 · 被引用 18 次
- Sampling-based inference for large linear models, with application to linearised LaplaceJavier Antorán, Shreyas Padhy, Riccardo Barbano, Eric T. Nalisnick 等ICLR 2023 · 被引用 1 次
相关 Paper
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate GradientsArtem Artemev, David R. Burt, Mark van der WilkICML 2021 · 被引用 28 次
- Scaling Gaussian Processes with Derivative Information Using Variational InferenceMisha Padidar, Xinran Zhu, Leo Huang, Jacob R. Gardner 等NeurIPS 2021 · 被引用 28 次
- Task-Agnostic Amortized Inference of Gaussian Process HyperparametersSulin Liu, Xingyuan Sun, Peter J. Ramadge, Ryan P. AdamsNeurIPS 2020 · 被引用 27 次
- Adjoint-aided inference of Gaussian process driven differential equationsPaterne Gahungu, Christopher W. Lanyon, Mauricio A. Álvarez, Engineer Bainomugisha 等NeurIPS 2022 · 被引用 6 次
- Preconditioning for Scalable Gaussian Process Hyperparameter OptimizationJonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham 等ICML 2022 · 被引用 36 次
