SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, Jiwei Li
摘要
While the self-attention mechanism has been widely used in a wide variety of tasks, it has the unfortunate property of a quadratic cost with respect to the input length, which makes it difficult to deal with long inputs. In this paper, we present a method for accelerating and structuring self-attentions: Sparse Adaptive Connection (SAC). In SAC, we regard the input sequence as a graph and attention operations are performed between linked nodes. In contrast with previous self-attention models with pre-defined structures (edges), the model learns to construct attention edges to improve task-specific performances. In this way, the model is able to select the most salient nodes and reduce the quadratic complexity regardless of the sequence length. Based on SAC, we show that previous variants of self-attention models are its special cases. Through extensive experiments on neural machine translation, language modeling, graph representation learning and image classification, we demonstrate SAC is competitive with state-of-the-art models while significantly reducing memory cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Hierarchical Graph Transformer with Adaptive Node SamplingZaixi Zhang, Qi Liu, Qingyong Hu, Chee-Kong LeeNeurIPS 2022 · 被引用 145 次
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat 等NeurIPS 2020 · 被引用 111 次
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu 等NeurIPS 2022 · 被引用 104 次
- GNN-LM: Language Modeling based on Global Contexts via GNNYuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun 等ICLR 2022 · 被引用 46 次
- Layer-wise Model Pruning based on Mutual InformationChun Fan, Jiwei Li, Tianwei Zhang, Xiang Ao 等EMNLP 2021
它引用的顶会 Paper7
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani 等ICCV 2019 · 被引用 1,149 次
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 被引用 555 次
- Adaptive Structural Fingerprints for Graph Attention NetworksKai Zhang, Yaokang Zhu, Jun Wang, Jie ZhangICLR 2020 · 被引用 87 次
相关 Paper
- Transformers meet Stochastic Block Models: Attention with Data-Adaptive Sparsity and CostSungjun Cho, Seonwoo Min, Jinwoo Kim, Moontae Lee 等NeurIPS 2022 · 被引用 5 次
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 被引用 11 次
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng 等DAC 2022 · 被引用 37 次
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen 等ASPLOS 2022 · 被引用 131 次
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 被引用 46 次
