Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian Processes
Jihao Andreas Lin, Shreyas Padhy, Bruno Mlodozeniec, Javier Antorán, José Miguel Hernández-Lobato
Abstract
Scaling hyperparameter optimisation to very large datasets remains an open problem in the Gaussian process community. This paper focuses on iterative methods, which use linear system solvers, like conjugate gradients, alternating projections or stochastic gradient descent, to construct an estimate of the marginal likelihood gradient. We discuss three key improvements which are applicable across solvers: (i) a pathwise gradient estimator, which reduces the required number of solver iterations and amortises the computational cost of making predictions, (ii) warm starting linear system solvers with the solution from the previous step, which leads to faster solver convergence at the cost of negligible bias, (iii) early stopping linear system solvers after a limited computational budget, which synergises with warm starting, allowing solver progress to accumulate over multiple marginal likelihood steps. These techniques provide speed-ups of up to when solving to tolerance, and decrease the average residual norm by up to when stopping early.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb8b6267-0555-4877-abf3-4a93ba43d222Cited by top-tier papers2
- Graph Random Features for Scalable Gaussian ProcessesMatthew Zhang, Jihao Andreas Lin, Krzysztof Choromanski, Adrian Weller et al.ICLR 2026 · 4 citations
- Scalable Gaussian Processes with Latent Kronecker StructureJihao Andreas Lin, Sebastian Ament, Maximilian Balandat, David Eriksson et al.ICML 2025
Builds on6
- Efficiently sampling functions from Gaussian process posteriorsJames T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky et al.ICML 2020 · 186 citations
- Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian ProcessesHao Chen, Lili Zheng, Raed Al Kontar, Garvesh RaskuttiNeurIPS 2020 · 49 citations
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz et al.NeurIPS 2023 · 34 citations
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate GradientsArtem Artemev, David R. Burt, Mark van der WilkICML 2021 · 28 citations
- Stochastic Gradient Descent for Gaussian Processes Done RightJihao Andreas Lin, Shreyas Padhy, Javier Antorán, Austin Tripp et al.ICLR 2024 · 17 citations
Related papers
- Preconditioning for Scalable Gaussian Process Hyperparameter OptimizationJonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham et al.ICML 2022 · 36 citations
- Bias-Free Scalable Gaussian Processes via Randomized TruncationsAndres Potapczynski, Luhuan Wu, Dan Biderman, Geoff Pleiss et al.ICML 2021 · 23 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- Turbocharging Gaussian Process Inference with Approximate Sketch-and-ProjectPratik Rathore, Zachary Frangella, Sachin Garg, Shaghayegh Fazliani et al.NeurIPS 2025 · 8 citations
- Probabilistic Unrolling: Scalable, Inverse-Free Maximum Likelihood Estimation for Latent Gaussian ModelsAlexander Lin, Bahareh Tolooshams, Yves F. Atchadé, Demba E. BaICML 2023 · 1 citation
