Calibrating Transformers via Sparse Gaussian Processes
Wenlong Chen, Yingzhen Li
Abstract
Transformer models have achieved profound success in prediction tasks in a wide range of applications in natural language processing, speech recognition and computer vision. Extending Transformer's success to safety-critical domains requires calibrated uncertainty estimation which remains under-explored. To address this, we propose Sparse Gaussian Process attention (SGPA), which performs Bayesian inference directly in the output space of multi-head attention blocks (MHAs) in transformer to calibrate its uncertainty. It replaces the scaled dot-product operation with a valid symmetric kernel and uses sparse Gaussian processes (SGP) techniques to approximate the posterior processes of MHA outputs. Empirically, on a suite of prediction tasks on text, images and graphs, SGPA-based Transformers achieve competitive predictive accuracy, while noticeably improving both in-distribution calibration and out-of-distribution robustness and detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 111 citations
- BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language ModelsYibin Wang, Haizhou Shi, Ligong Han, Dimitris N. Metaxas et al.NeurIPS 2024 · 63 citations
- Elliptical AttentionStefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan NguyenNeurIPS 2024 · 12 citations
- Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component AnalysisRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 11 citations
- Variational Uncertainty Decomposition for In-Context LearningI. Shavindra Jayasekera, Jacob Si, Filippo Valdettaro, Wenlong Chen et al.NeurIPS 2025 · 7 citations
Related papers
- Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian ProcessesYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensICML 2024 · 4 citations
- SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep LearningWenyuan Zhao, Rui Tuo, Chao TianICML 2026
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior AdaptationGeorge Whittle, Juliusz Ziomek, Jacob Rawling, Michael A OsborneICML 2026 · 13 citations
- GPan-LoRA: Gaussian Process Amortized Networks for Bayesian Low-Rank Adaptation in Large Language ModelsWeifeng Zhang, Wenyuan Zhao, Amir Hossein Rahmati, Yucheng Wang et al.ICML 2026
- Enabling fast uncertainty estimation: accelerating bayesian transformers via algorithmic and hardware optimizationsHongxiang Fan, Martin Ferianc, Wayne LukDAC 2022 · 7 citations
