Bayesian Attention Belief Networks
Shujian Zhang, Xinjie Fan, Bo Chen, Mingyuan Zhou
Abstract
Attention-based neural networks have achieved state-of-the-art results on a wide range of tasks. Most such models use deterministic attention while stochastic attention is less explored due to the optimization difficulties or complicated model design. This paper introduces Bayesian attention belief networks, which construct a decoder network by modeling unnormalized attention weights with a hierarchy of gamma distributions, and an encoder network by stacking Weibull distributions with a deterministic-upward-stochastic-downward structure to approximate the posterior. The resulting auto-encoding networks can be optimized in a differentiable way with a variational lower bound. It is simple to convert any models with deterministic attention, including pretrained ones, to the proposed Bayesian attention belief networks. On a variety of language understanding tasks, we show that our method outperforms deterministic attention and state-of-the-art stochastic attention in accuracy, uncertainty estimation, generalization across domains, and robustness to adversarial attacks. We further demonstrate the general applicability of our method on neural machine translation and visual question answering, showing great potential of incorporating our method into various attention-related tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64a11e67-471f-4dec-be82-82ca3daad31bCited by top-tier papers16
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 111 citations
- BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language ModelsYibin Wang, Haizhou Shi, Ligong Han, Dimitris N. Metaxas et al.NeurIPS 2024 · 63 citations
- POUF: Prompt-Oriented Unsupervised Fine-tuning for Large Pre-trained ModelsKorawat Tanwisuth, Shujian Zhang, Huangjie Zheng, Pengcheng He et al.ICML 2023 · 44 citations
- Uncertainty-Guided Probabilistic Transformer for Complex Action RecognitionHongji Guo, Hanjing Wang, Qiang JiCVPR 2022 · 42 citations
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng et al.NeurIPS 2023 · 39 citations
Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Bayesian Attention ModulesXinjie Fan, Shujian Zhang, Bo Chen, Mingyuan ZhouNeurIPS 2020 · 78 citations
Related papers
- Alignment Attention by Matching Key and Query DistributionsShujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth et al.NeurIPS 2021 · 20 citations
- A Bit More Bayesian: Domain-Invariant Learning with UncertaintyZehao Xiao, Jiayi Shen, Xiantong Zhen, Ling Shao et al.ICML 2021 · 47 citations
- Regularizing Attention Networks for Anomaly Detection in Visual Question AnsweringDoyup Lee, Yeongjae Cheon, Wook-Shin HanAAAI 2021 · 17 citations
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 5 citations
- Transformer Uncertainty Estimation with Hierarchical Stochastic AttentionJiahuan Pei, Cheng Wang, György SzarvasAAAI 2022 · 33 citations
