Scalable First-Order Bayesian Optimization via Structured Automatic Differentiation
Sebastian E. Ament, Carla P. Gomes
Abstract
Bayesian Optimization (BO) has shown great promise for the global optimization of functions that are expensive to evaluate, but despite many successes, standard approaches can struggle in high dimensions. To improve the performance of BO, prior work suggested incorporating gradient information into a Gaussian process surrogate of the objective, giving rise to kernel matrices of size for observations in dimensions. Naïvely multiplying with (resp. inverting) these matrices requires (resp. )) operations, which becomes infeasible for moderate dimensions and sample sizes. Here, we observe that a wide range of kernels gives rise to structured matrices, enabling an exact matrix-vector multiply for gradient observations and for Hessian observations. Beyond canonical kernel classes, we derive a programmatic approach to leveraging this type of structure for transformations and combinations of the discussed kernel classes, which constitutes a structure-aware automatic differentiation algorithm. Our methods apply to virtually all canonical kernels and automatically extend to complex kernels, like the neural network, radial basis function network, and spectral mixture kernels without any additional derivations, enabling flexible, problem-dependent modeling while scaling first-order BO to high .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Unexpected Improvements to Expected Improvement for Bayesian OptimizationSebastian Ament, Samuel Daulton, David Eriksson, Maximilian Balandat et al.NeurIPS 2023 · 280 citations
- SurCo: Learning Linear SURrogates for COmbinatorial Nonlinear Optimization ProblemsAaron M. Ferber, Taoan Huang, Daochen Zha, Martin Schubert et al.ICML 2023 · 25 citations
- BayeSQP: Bayesian Optimization through Sequential Quadratic ProgrammingPaul Brunzema, Sebastian TrimpeNeurIPS 2025 · 7 citations
- Auto-Differentiation of Relational Computations for Very Large Scale Machine LearningYuxin Tang, Zhimin Ding, Dimitrije Jankov, Binhang Yuan et al.ICML 2023 · 7 citations
- Optimizing the Unknown: Black Box Bayesian Optimization with Energy-Based Model and Reinforcement LearningRuiyao Miao, Junren Xiao, Shiya Tsang, Hui Xiong et al.NeurIPS 2025 · 2 citations
Builds on5
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton et al.NeurIPS 2020 · 686 citations
- Scaling Gaussian Processes with Derivative Information Using Variational InferenceMisha Padidar, Xinran Zhu, Leo Huang, Jacob R. Gardner et al.NeurIPS 2021 · 28 citations
- High-Dimensional Gaussian Process Inference with DerivativesFilip de Roos, Alexandra Gessner, Philipp HennigICML 2021 · 24 citations
- Sparse Bayesian Learning via Stepwise RegressionSebastian E. Ament, Carla P. GomesICML 2021 · 11 citations
- Sequential Bayesian Experimental Design with Variable Cost StructureSue Zheng, David S. Hayden, Jason Pacheco, John W. Fisher IIINeurIPS 2020 · 10 citations
Related papers
- Modeling All Response Surfaces in One for Conditional Search SpacesJiaxing Li, Wei Liu, Chao Xue, Yibing Zhan et al.AAAI 2025 · 1 citation
- Standard Gaussian Process is All You Need for High-Dimensional Bayesian OptimizationZhitong Xu, Haitao Wang, Jeff M. Phillips, Shandian ZheICLR 2025
- From Sorting Algorithms to Scalable Kernels: Bayesian Optimization in High-Dimensional Permutation SpacesZikai Xie, Linjiang ChenICLR 2026 · 2 citations
- A Study of Bayesian Neural Network Surrogates for Bayesian OptimizationYucen Lily Li, Tim G. J. Rudner, Andrew Gordon WilsonICLR 2024 · 59 citations
- Sample-Then-Optimize Batch Neural Thompson SamplingZhongxiang Dai, Yao Shu, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2022 · 33 citations
