Calibrating Transformers via Sparse Gaussian Processes
Wenlong Chen, Yingzhen Li
摘要
Transformer models have achieved profound success in prediction tasks in a wide range of applications in natural language processing, speech recognition and computer vision. Extending Transformer's success to safety-critical domains requires calibrated uncertainty estimation which remains under-explored. To address this, we propose Sparse Gaussian Process attention (SGPA), which performs Bayesian inference directly in the output space of multi-head attention blocks (MHAs) in transformer to calibrate its uncertainty. It replaces the scaled dot-product operation with a valid symmetric kernel and uses sparse Gaussian processes (SGP) techniques to approximate the posterior processes of MHA outputs. Empirically, on a suite of prediction tasks on text, images and graphs, SGPA-based Transformers achieve competitive predictive accuracy, while noticeably improving both in-distribution calibration and out-of-distribution robustness and detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 被引用 111 次
- BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language ModelsYibin Wang, Haizhou Shi, Ligong Han, Dimitris N. Metaxas 等NeurIPS 2024 · 被引用 63 次
- Elliptical AttentionStefan K. Nielsen, Laziz U. Abdullaev, Rachel S. Y. Teo, Tan NguyenNeurIPS 2024 · 被引用 12 次
- Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component AnalysisRachel S. Y. Teo, Tan M. NguyenNeurIPS 2024 · 被引用 11 次
- Variational Uncertainty Decomposition for In-Context LearningI. Shavindra Jayasekera, Jacob Si, Filippo Valdettaro, Wenlong Chen 等NeurIPS 2025 · 被引用 7 次
相关 Paper
- Self-Attention through Kernel-Eigen Pair Sparse Variational Gaussian ProcessesYingyi Chen, Qinghua Tao, Francesco Tonin, Johan A. K. SuykensICML 2024 · 被引用 4 次
- SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep LearningWenyuan Zhao, Rui Tuo, Chao TianICML 2026
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior AdaptationGeorge Whittle, Juliusz Ziomek, Jacob Rawling, Michael A OsborneICML 2026 · 被引用 13 次
- GPan-LoRA: Gaussian Process Amortized Networks for Bayesian Low-Rank Adaptation in Large Language ModelsWeifeng Zhang, Wenyuan Zhao, Amir Hossein Rahmati, Yucheng Wang 等ICML 2026
- Enabling fast uncertainty estimation: accelerating bayesian transformers via algorithmic and hardware optimizationsHongxiang Fan, Martin Ferianc, Wayne LukDAC 2022 · 被引用 7 次
