SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
Xiaoya Li, Yuxian Meng, Mingxin Zhou, Qinghong Han, Fei Wu, Jiwei Li
Abstract
While the self-attention mechanism has been widely used in a wide variety of tasks, it has the unfortunate property of a quadratic cost with respect to the input length, which makes it difficult to deal with long inputs. In this paper, we present a method for accelerating and structuring self-attentions: Sparse Adaptive Connection (SAC). In SAC, we regard the input sequence as a graph and attention operations are performed between linked nodes. In contrast with previous self-attention models with pre-defined structures (edges), the model learns to construct attention edges to improve task-specific performances. In this way, the model is able to select the most salient nodes and reduce the quadratic complexity regardless of the sequence length. Based on SAC, we show that previous variants of self-attention models are its special cases. Through extensive experiments on neural machine translation, language modeling, graph representation learning and image classification, we demonstrate SAC is competitive with state-of-the-art models while significantly reducing memory cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Hierarchical Graph Transformer with Adaptive Node SamplingZaixi Zhang, Qi Liu, Qingyong Hu, Chee-Kong LeeNeurIPS 2022 · 145 citations
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat et al.NeurIPS 2020 · 111 citations
- Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationBotao Yu, Peiling Lu, Rui Wang, Wei Hu et al.NeurIPS 2022 · 104 citations
- GNN-LM: Language Modeling based on Global Contexts via GNNYuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun et al.ICLR 2022 · 46 citations
- Layer-wise Model Pruning based on Mutual InformationChun Fan, Jiwei Li, Tianwei Zhang, Xiang Ao et al.EMNLP 2021
Builds on7
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- Adaptive Structural Fingerprints for Graph Attention NetworksKai Zhang, Yaokang Zhu, Jun Wang, Jie ZhangICLR 2020 · 87 citations
Related papers
- Transformers meet Stochastic Block Models: Attention with Data-Adaptive Sparsity and CostSungjun Cho, Seonwoo Min, Jinwoo Kim, Moontae Lee et al.NeurIPS 2022 · 5 citations
- Long-range Sequence Modeling with Predictable Sparse AttentionYimeng Zhuang, Jing Zhang, Mei TuACL 2022 · 11 citations
- SALO: an efficient spatial accelerator enabling hybrid sparse attention mechanisms for long sequencesGuan Shen, Jieru Zhao, Quan Chen, Jingwen Leng et al.DAC 2022 · 37 citations
- DOTA: detect and omit weak attentions for scalable transformer accelerationZheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen et al.ASPLOS 2022 · 131 citations
- AutoAttend: Automated Attention Representation SearchChaoyu Guan, Xin Wang, Wenwu ZhuICML 2021 · 46 citations
