Kernel Identification Through Transformers
Fergus Simpson, Ian Davies, Vidhi Lalchand, Alessandro Vullo, Nicolas Durrande, Carl Edward Rasmussen
Abstract
Kernel selection plays a central role in determining the performance of Gaussian Process (GP) models, as the chosen kernel determines both the inductive biases and prior support of functions under the GP prior. This work addresses the challenge of constructing custom kernel functions for high-dimensional GP regression models. Drawing inspiration from recent progress in deep learning, we introduce a novel approach named KITT: Kernel Identification Through Transformers. KITT exploits a transformer-based architecture to generate kernel recommendations in under 0.1 seconds, which is several orders of magnitude faster than conventional kernel search algorithms. We train our model using synthetic data generated from priors over a vocabulary of known kernels. By exploiting the nature of the selfattention mechanism, KITT is able to process datasets with inputs of arbitrary dimension. We demonstrate that kernels chosen by KITT yield strong performance over a diverse collection of regression benchmarks. * Work undertaken while at Secondmind 1 also called covariance function or covariance kernel 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b290439-9fcf-4a29-967d-276518cedcb5Cited by top-tier papers5
- Neural Diffusion ProcessesVincent Dutordoir, Alan Saul, Zoubin Ghahramani, Fergus SimpsonICML 2023 · 52 citations
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMsRichard Cornelius Suwandi, Feng Yin, Juntao Wang, Renjie Li et al.NeurIPS 2025 · 19 citations
- Practical Equivariances via Relational Conditional Neural ProcessesDaolang Huang, Manuel Haussmann, Ulpu Remes, S. T. John et al.NeurIPS 2023 · 14 citations
- On the Identifiability and Interpretability of Gaussian Process ModelsJiawen Chen, Wancen Mu, Yun Li, Didong LiNeurIPS 2023 · 8 citations
- EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian OptimizationMujin Cheon, Jay H. Lee, Dong-Yeun Koh, Calvin TsayICML 2025
Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Efficient Transformers in Reinforcement Learning using Actor-Learner DistillationEmilio Parisotto, Ruslan SalakhutdinovICLR 2021 · 51 citations
- Task-Agnostic Amortized Inference of Gaussian Process HyperparametersSulin Liu, Xingyuan Sun, Peter J. Ramadge, Ryan P. AdamsNeurIPS 2020 · 27 citations
- Meshed-Memory Transformer for Image CaptioningMarcella Cornia, Matteo Stefanini, Lorenzo Baraldi, Rita CucchiaraCVPR 2020
Related papers
- Structural Kernel Search via Bayesian Optimization and Symbolical Optimal TransportMatthias Bitzer, Mona Meister, Christoph ZimmerNeurIPS 2022 · 11 citations
- Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel MethodsZhaiming Shen, Alexander Hsu, Rongjie Lai, Wenjing LiaoICLR 2026 · 13 citations
- Graph Neural Network-Inspired Kernels for Gaussian Processes in Semi-Supervised LearningZehao Niu, Mihai Anitescu, Jie ChenICLR 2023 · 1 citation
- Kernel Functional OptimisationArun Kumar Anjanapura Venkatesh, Alistair Shilton, Santu Rana, Sunil Gupta et al.NeurIPS 2021 · 6 citations
- Longitudinal Deep Kernel Gaussian Process RegressionJunjie Liang, Yanting Wu, Dongkuan Xu, Vasant G. HonavarAAAI 2021 · 9 citations
