Preconditioning for Scalable Gaussian Process Hyperparameter Optimization
Jonathan Wenger, Geoff Pleiss, Philipp Hennig, John P. Cunningham, Jacob R. Gardner
Abstract
Gaussian process hyperparameter optimization requires linear solves with, and log -determinants of, large kernel matrices. Iterative numerical tech-niques are becoming popular to scale to larger datasets, relying on the conjugate gradient method (CG) for the linear solves and stochastic trace estimation for the log -determinant. This work introduces new algorithmic and theoretical in-sights for preconditioning these computations. While preconditioning is well understood in the context of CG, we demonstrate that it can also accelerate convergence and reduce variance of the estimates for the log -determinant and its derivative. We prove general probabilistic error bounds for the preconditioned computation of the log -determinant, log -marginal likelihood and its derivatives. Additionally, we derive specific rates for a range of kernel-preconditioner combinations, showing that up to exponential convergence can be achieved. Our theoretical results enable prov-ably efficient optimization of kernel hyperparameters, which we validate empirically on large-scale benchmark problems. There our approach accelerates training by up to an order of magnitude.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b98ab40-fb53-4975-a3be-831fff57e73eCited by top-tier papers6
- Posterior and Computational Uncertainty in Gaussian ProcessesJonathan Wenger, Geoff Pleiss, Marvin Pförtner, Philipp Hennig et al.NeurIPS 2022 · 31 citations
- Computation-Aware Gaussian Processes: Model Selection And Linear-Time InferenceJonathan Wenger, Kaiwen Wu, Philipp Hennig, Jacob R. Gardner et al.NeurIPS 2024 · 15 citations
- Implicit Manifold Gaussian Process RegressionBernardo Fichera, Slava Borovitskiy, Andreas Krause, Aude Gemma BillardNeurIPS 2023 · 10 citations
- Gradients of Functions of Large MatricesNicholas Krämer, Pablo Moreno-Muñoz, Hrittik Roy, Søren HaubergNeurIPS 2024 · 6 citations
- Probabilistic Unrolling: Scalable, Inverse-Free Maximum Likelihood Estimation for Latent Gaussian ModelsAlexander Lin, Bahareh Tolooshams, Yves F. Atchadé, Demba E. BaICML 2023 · 1 citation
Builds on3
- Optimal Sketching for Trace EstimationShuli Jiang, Hai Pham, David P. Woodruff, Qiuyi (Richard) ZhangNeurIPS 2021 · 28 citations
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate GradientsArtem Artemev, David R. Burt, Mark van der WilkICML 2021 · 28 citations
- Bias-Free Scalable Gaussian Processes via Randomized TruncationsAndres Potapczynski, Luhuan Wu, Dan Biderman, Geoff Pleiss et al.ICML 2021 · 23 citations
Related papers
- Improving Linear System Solvers for Hyperparameter Optimisation in Iterative Gaussian ProcessesJihao Andreas Lin, Shreyas Padhy, Bruno Mlodozeniec, Javier Antorán et al.NeurIPS 2024 · 6 citations
- Turbocharging Gaussian Process Inference with Approximate Sketch-and-ProjectPratik Rathore, Zachary Frangella, Sachin Garg, Shaghayegh Fazliani et al.NeurIPS 2025 · 8 citations
- Kernel Methods Through the Roof: Handling Billions of Points EfficientlyGiacomo Meanti, Luigi Carratino, Lorenzo Rosasco, Alessandro RudiNeurIPS 2020 · 138 citations
- KernelMatmul: Scaling Gaussian Processes to Large Time SeriesTilman Hoffbauer, Holger H. Hoos, Jakob BossekAAAI 2025
- gp2Scale: A Class of Compactly Supported Non-Stationary Kernels and Distributed Computing for Exact Gaussian Processes on 10 Million Data PointsMarcus Noack, Mark Risser, HENGRUI LUO, Vardaan Tekriwal et al.ICML 2026 · 1 citation
