Bayesian Attention Modules
Xinjie Fan, Shujian Zhang, Bo Chen, Mingyuan Zhou
Abstract
Attention modules, as simple and effective tools, have not only enabled deep neural networks to achieve state-of-the-art results in many domains, but also enhanced their interpretability. Most current models use deterministic attention modules due to their simplicity and ease of optimization. Stochastic counterparts, on the other hand, are less popular despite their potential benefits. The main reason is that stochastic attention often introduces optimization issues or requires significant model changes. In this paper, we propose a scalable stochastic version of attention that is easy to implement and optimize. We construct simplex-constrained attention distributions by normalizing reparameterizable distributions, making the training process differentiable. We learn their parameters in a Bayesian framework where a data-dependent prior is introduced for regularization. We apply the proposed stochastic attention modules to various attention-based models, with applications to graph node classification, visual question answering, image captioning, machine translation, and language understanding. Our experiments show the proposed method brings consistent improvements over the corresponding baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b1fbe8a-fb01-4290-8a4f-cfedd9186f61Cited by top-tier papers27
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 111 citations
- BLoB: Bayesian Low-Rank Adaptation by Backpropagation for Large Language ModelsYibin Wang, Haizhou Shi, Ligong Han, Dimitris N. Metaxas et al.NeurIPS 2024 · 63 citations
- Sawtooth Factorial Topic Embeddings Guided Gamma Belief NetworkZhibin Duan, Dongsheng Wang, Bo Chen, Chaojie Wang et al.ICML 2021 · 49 citations
- POUF: Prompt-Oriented Unsupervised Fine-tuning for Large Pre-trained ModelsKorawat Tanwisuth, Shujian Zhang, Huangjie Zheng, Pengcheng He et al.ICML 2023 · 44 citations
- Preference-grounded Token-level Guidance for Language Model Fine-tuningShentao Yang, Shujian Zhang, Congying Xia, Yihao Feng et al.NeurIPS 2023 · 39 citations
Builds on3
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Adaptive Correlated Monte Carlo for Contextual Categorical Sequence GenerationXinjie Fan, Yizhe Zhang, Zhendong Wang, Mingyuan ZhouICLR 2020 · 4 citations
Related papers
- Bayesian Attention Belief NetworksShujian Zhang, Xinjie Fan, Bo Chen, Mingyuan ZhouICML 2021 · 38 citations
- Alignment Attention by Matching Key and Query DistributionsShujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth et al.NeurIPS 2021 · 20 citations
- Contextual Dropout: An Efficient Sample-Dependent Dropout ModuleXinjie Fan, Shujian Zhang, Korawat Tanwisuth, Xiaoning Qian et al.ICLR 2021 · 34 citations
- U-CAM: Visual Explanation Using Uncertainty Based Class Activation MapsBadri N. Patro, Mayank Lunayach, Shivansh Patel, Vinay P. NamboodiriICCV 2019 · 82 citations
- Affine-Scaled Attention: Towards Flexible and Stable Transformer AttentionJeongin Bae, baeseong park, Gunho Park, Minsub Kim et al.ICML 2026 · 1 citation
