Not All Attention Is Needed: Gated Attention Network for Sequence Data
Lanqing Xue, Xiaopeng Li, Nevin L. Zhang
摘要
Although deep neural networks generally have fixed network structures, the concept of dynamic mechanism has drawn more and more attention in recent years. Attention mechanisms compute input-dependent dynamic attention weights for aggregating a sequence of hidden states. Dynamic network configuration in convolutional neural networks (CNNs) selectively activates only part of the network at a time for different inputs. In this paper, we combine the two dynamic mechanisms for text classification tasks. Traditional attention mechanisms attend to the whole sequence of hidden states for an input sentence, while in most cases not all attention is needed especially for long sequences. We propose a novel method called Gated Attention Network (GA-Net) to dynamically select a subset of elements to attend to using an auxiliary network, and compute attention weights to aggregate the selected elements. It avoids a significant amount of unnecessary computation on unattended elements, and allows the model to pay attention to important parts of the sequence. Experiments in various datasets show that the proposed method achieves better performance compared with all baseline models with global or local attention while requiring less computation and achieving better interpretability. It is also promising to extend the idea to more complex attention-based models, such as transformers and seq-to-seq models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- Generating Diversified Comments via Reader-Aware Topic Modeling and Saliency DetectionWei Wang, Piji Li, Hai-Tao ZhengAAAI 2021 · 被引用 16 次
- Embodied CoT Distillation From LLM To Off-the-shelf AgentsWonje Choi, Woo Kyung Kim, Minjong Yoo, Honguk WooICML 2024 · 被引用 13 次
相关 Paper
- ACT: an Attentive Convolutional Transformer for Efficient Text ClassificationPengfei Li, Peixiang Zhong, Kezhi Mao, Dongzhe Wang 等AAAI 2021 · 被引用 47 次
- Merging Statistical Feature via Adaptive Gate for Improved Text ClassificationXianming Li, Zongxi Li, Haoran Xie, Qing LiAAAI 2021 · 被引用 55 次
- Implicit Kernel AttentionKyungwoo Song, Yohan Jung, Dongjun Kim, Il-Chul MoonAAAI 2021 · 被引用 18 次
- One-shot Graph Neural Architecture Search with Dynamic Search SpaceYanxi Li, Zean Wen, Yunhe Wang, Chang XuAAAI 2021 · 被引用 54 次
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz 等ICLR 2020 · 被引用 152 次
