Transformer Uncertainty Estimation with Hierarchical Stochastic Attention
Jiahuan Pei, Cheng Wang, György Szarvas
摘要
Transformers are state-of-the-art in a wide range of NLP tasks and have also been applied to many real-world products. Understanding the reliability and certainty of transformer model predictions is crucial for building trustable machine learning applications, e.g., medical diagnosis. Although many recent transformer extensions have been proposed, the study of the uncertainty estimation of transformer models is under-explored. In this work, we propose a novel way to enable transformers to have the capability of uncertainty estimation and, meanwhile, retain the original predictive performance. This is achieved by learning a hierarchical stochastic self-attention that attends to values and a set of learnable centroids, respectively. Then new attention heads are formed with a mixture of sampled centroids using the Gumbel-Softmax trick. We theoretically show that the selfattention approximation by sampling from a Gumbel distribution is upper bounded. We empirically evaluate our model on two text classification tasks with both in-domain (ID) and outof-domain (OOD) datasets. The experimental results demonstrate that our approach: (1) achieves the best predictive performance and uncertainty trade-off among compared methods; (2) exhibits very competitive (in most cases, improved) predictive performance on ID datasets; (3) is on par with Monte Carlo dropout and ensemble methods in uncertainty estimation on OOD datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Energy-Based Transformers are Scalable Learners and ThinkersAlexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han 等ICLR 2026 · 被引用 38 次
- Probabilistic Transformer: Modelling Ambiguities and Distributions for RNA Folding and Molecule DesignJörg K. H. Franke, Frederic Runge, Frank HutterNeurIPS 2022 · 被引用 19 次
- On the Calibration and Uncertainty with Pólya-Gamma Augmentation for Dialog Retrieval ModelsTong Ye, Shijing Si, Jianzong Wang, Ning Cheng 等AAAI 2023 · 被引用 3 次
- Uncertainty-Instructed Structure Injection for Generalizable HD Map ConstructionXiaolu Liu, Ruizi Yang, Song Wang, Wentong Li 等CVPR 2025
- S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object TrackingTao Tang, Lijun Zhou, Pengkun Hao, Zihang He 等ICML 2025
它引用的顶会 Paper6
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep LearningArsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, Dmitry P. VetrovICLR 2020 · 被引用 354 次
- Fast Transformers with Clustered AttentionApoorv Vyas, Angelos Katharopoulos, François FleuretNeurIPS 2020 · 被引用 193 次
- On Position Embeddings in BERTBenyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 等ICLR 2021 · 被引用 129 次
- Doubly Stochastic Variational Inference for Neural Processes with Hierarchical Latent VariablesQi Wang, Herke van HoofICML 2020 · 被引用 50 次
相关 Paper
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 被引用 182 次
- Uncertainty Estimation of Transformer Predictions for Misclassification DetectionArtem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun 等ACL 2022 · 被引用 59 次
- Calibrating Transformers via Sparse Gaussian ProcessesWenlong Chen, Yingzhen LiICLR 2023
- Estimating Epistemic and Aleatoric Uncertainty with a Single ModelMatthew Chan, Maria Molina, Chris MetzlerNeurIPS 2024 · 被引用 76 次
- How Does Selective Mechanism Improve Self-Attention Networks?Xinwei Geng, Longyue Wang, Xing Wang, Bing Qin 等ACL 2020 · 被引用 30 次
