Transformer Uncertainty Estimation with Hierarchical Stochastic Attention
Jiahuan Pei, Cheng Wang, György Szarvas
Abstract
Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine learning applications, e.g., medical diagnosis. Although many recent transformer extensions have been proposed, the study of the uncertainty estimation of transformer models is under-explored. In this work, we propose a novel way to enable transformers to have the capability of uncertainty estimation and, meanwhile, retain the original predictive performance. This is achieved by learning a hierarchical stochastic self-attention that attends to values and a set of learnable centroids, respectively. Then new attention heads are formed with a mixture of sampled centroids using the Gumbel-Softmax trick. We theoretically show that the selfattention approximation by sampling from a Gumbel distribution is upper bounded. We empirically evaluate our model on two text classification tasks with both in-domain (ID) and outof-domain (OOD) datasets. The experimental results demonstrate that our approach: (1) achieves the best predictive performance and uncertainty trade-off among compared methods; (2) exhibits very competitive (in most cases, improved) predictive performance on ID datasets; (3) is on par with Monte Carlo dropout and ensemble methods in uncertainty estimation on OOD datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b81ea120-2aa5-436f-80ce-9bf824af60f6Cited by top-tier papers6
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han et al.ICLR 2026 · 38 citations
- Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule DesignJörg K. H. Franke, Frederic Runge, Frank HutterNeurIPS 2022 · 19 citations
- On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsTong Ye, Shijing Si, Jianzong Wang, Ning Cheng et al.AAAI 2023 · 3 citations
- Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionXiaolu Liu, Ruizi Yang, Song Wang, Wentong Li et al.CVPR 2025
- S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object TrackingTao Tang, Lijun Zhou, Pengkun Hao, Zihang He et al.ICML 2025
Builds on6
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 354 citations
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 193 citations
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang et al.ICLR 2021 · 129 citations
- Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent VariablesQi Wang, Herke van HoofICML 2020 · 50 citations
Related papers
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Uncertainty Estimation of Transformer Predictions for Misclassification DetectionArtem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun et al.ACL 2022 · 59 citations
- Calibrating Transformers via Sparse Gaussian ProcessesWenlong Chen, Yingzhen LiICLR 2023
- Estimating Epistemic and Aleatoric Uncertainty with a Single ModelMatthew Chan, Maria Molina, Chris MetzlerNeurIPS 2024 · 76 citations
- How Does Selective Mechanism Improve Self-Attention Networks?Xinwei Geng, Longyue Wang, Xing Wang, Bing Qin et al.ACL 2020 · 30 citations
