High-Dimensional Gaussian Process Inference with Derivatives
Filip de Roos, Alexandra Gessner, Philipp Hennig
Abstract
Although it is widely known that Gaussian processes can be conditioned on observations of the gradient, this functionality is of limited use due to the prohibitive computational cost of in data points and dimension . The dilemma of gradient observations is that a single one of them comes at the same cost as independent function evaluations, so the latter are often preferred. Careful scrutiny reveals, however, that derivative observations give rise to highly structured kernel Gram matrices for very general classes of kernels (inter alia, stationary kernels). We show that in the low-data regime , the Gram matrix can be decomposed in a manner that reduces the cost of inference to (i.e., linear in the number of dimensions) and, in special cases, to . This reduction in complexity opens up new use-cases for inference with gradients especially in the high-dimensional regime, where the information-to-cost ratio of gradient observations significantly increases. We demonstrate this potential in a variety of tasks relevant for machine learning, such as optimization and Hamiltonian Monte Carlo with predictive gradients.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caf38627-afa5-4584-b441-3bd506219d69Cited by top-tier papers6
- Unexpected Improvements to Expected Improvement for Bayesian OptimizationSebastian Ament, Samuel Daulton, David Eriksson, Maximilian Balandat et al.NeurIPS 2023 · 280 citations
- Scaling Gaussian Processes with Derivative Information Using Variational InferenceMisha Padidar, Xinran Zhu, Leo Huang, Jacob R. Gardner et al.NeurIPS 2021 · 28 citations
- Scalable First-Order Bayesian Optimization via Structured Automatic DifferentiationSebastian E. Ament, Carla P. GomesICML 2022 · 12 citations
- Monotonicity and Double Descent in Uncertainty Estimation with Gaussian ProcessesLiam Hodgkinson, Christopher van der Heide, Fred Roosta, Michael W. MahoneyICML 2023 · 9 citations
- BayeSQP: Bayesian Optimization through Sequential Quadratic ProgrammingPaul Brunzema, Sebastian TrimpeNeurIPS 2025 · 7 citations
Builds on2
Related papers
- Scalable Gaussian Processes with Latent Kronecker StructureJihao Andreas Lin, Sebastian Ament, Maximilian Balandat, David Eriksson et al.ICML 2025
- The Price of Linear Time: Error Analysis of Structured Kernel InterpolationAlexander Moreno, Justin Xiao, Jonathan MeiICML 2025
- Sampling from Gaussian Process Posteriors using Stochastic Gradient DescentJihao Andreas Lin, Javier Antorán, Shreyas Padhy, David Janz et al.NeurIPS 2023 · 34 citations
- KernelMatmul: Scaling Gaussian Processes to Large Time SeriesTilman Hoffbauer, Holger H. Hoos, Jakob BossekAAAI 2025
- Variational Sparse Inverse Cholesky Approximation for Latent Gaussian Processes via Double Kullback-Leibler MinimizationJian Cao, Myeongjong Kang, Felix Jimenez, Huiyan Sang et al.ICML 2023 · 12 citations
