Amortized Variational Deep Kernel Learning
Alan L. S. Matias, César Lincoln C. Mattos, João Paulo Pordeus Gomes, Diego Mesquita
Abstract
Deep kernel learning (DKL) marries the uncertainty quantification of Gaussian processes (GPs) and the representational power of deep neural networks. However, training DKL is challenging and often leads to overfitting. Most notably, DKL often learns "non-local" kernels -incurring spurious correlations. To remedy this issue, we propose using amortized inducing points and a parameter-sharing scheme, which ties together the amortization and DKL networks. This design imposes an explicit dependency between the ELBO's model fit and capacity terms. In turn, this prevents the former from dominating the optimization procedure and incurring the aforementioned spurious correlations. Extensive experiments show that our resulting method, amortized varitional DKL (AVDKL), i) consistently outperforms DKL and standard GPs for tabular data; ii) achieves significantly higher accuracy than DKL in node classification tasks; and iii) leads to substantially better accuracy and negative loglikelihood than DKL on CIFAR100.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9de0465d-420a-4f92-ad43-31e5c4942285Cited by top-tier papers2
- Transformed Latent Variable Multi-Output Gaussian ProcessesXiaoyu Jiang, Xinxing Shi, Sokratia Georgaka, Magnus Rattray et al.ICML 2026
- GPan-LoRA: Gaussian Process Amortized Networks for Bayesian Low-Rank Adaptation in Large Language ModelsWeifeng Zhang, Wenyuan Zhao, Amir Hossein Rahmati, Yucheng Wang et al.ICML 2026
Builds on8
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee et al.ICLR 2020 · 254 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Personalized Federated Learning With Gaussian ProcessesIdan Achituve, Aviv Shamsian, Aviv Navon, Gal Chechik et al.NeurIPS 2021 · 137 citations
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 72 citations
Related papers
- Longitudinal Deep Kernel Gaussian Process RegressionJunjie Liang, Yanting Wu, Dongkuan Xu, Vasant G. HonavarAAAI 2021 · 9 citations
- Input Dependent Sparse Gaussian ProcessesBahram Jafrasteh, Carlos Villacampa-Calvo, Daniel Hernández-LobatoICML 2022 · 7 citations
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson et al.ICML 2023 · 15 citations
- Fully Bayesian Autoencoders with Latent Sparse Gaussian ProcessesBa-Hien Tran, Babak Shahbaba, Stephan Mandt, Maurizio FilipponeICML 2023 · 9 citations
