The Information Pathways Hypothesis: Transformers are Dynamic Self-Ensembles
Md. Shamim Hussain, Mohammed J. Zaki, Dharmashankar Subramanian
摘要
Transformers use the dense self-attention mechanism which gives a lot of flexibility for long-range connectivity. Over multiple layers of a deep transformer, the number of possible connectivity patterns increases exponentially. However, very few of these contribute to the performance of the network, and even fewer are essential. We hypothesize that there are sparsely connected sub-networks within a transformer, called information pathways which can be trained independently. However, the dynamic (i.e., input-dependent) nature of these pathways makes it difficult to prune dense self-attention during training. But the overall distribution of these pathways is often predictable. We take advantage of this fact to propose Stochastically Subsampled self-Attention (SSA) -a general-purpose training strategy for transformers that can reduce both the memory and computational cost of self-attention by 4 to 8 times during training while also serving as a regularization method -improving generalization over dense training. We show that an ensemble of sub-models can be formed from the subsampled pathways within a network, which can achieve better performance than its densely attended counterpart. We perform experiments on a variety of NLP, computer vision and graph learning tasks in both generative and discriminative settings to provide empirical evidence for our claims and show the effectiveness of the proposed method. CCS CONCEPTS • Computing methodologies → Neural networks; Ensemble methods; Artificial intelligence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Triplet Interaction Improves Graph Transformers: Accurate Molecular Graph Learning with Triplet Graph TransformersMd. Shamim Hussain, Mohammed J. Zaki, Dharmashankar SubramanianICML 2024 · 被引用 19 次
- LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language ModelsLong Chen, Xiaotian Song, Yanan SunAAAI 2026
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
相关 Paper
- Revisiting Vision Transformer from the View of Path EnsembleShuning Chang, Pichao Wang, Hao Luo, Fan Wang 等ICCV 2023 · 被引用 8 次
- Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersLorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim SompolinskyNeurIPS 2024 · 被引用 12 次
- Adaptive Depth Networks with Skippable Sub-PathsWoochul Kang, Hyungseop LeeNeurIPS 2024 · 被引用 5 次
- Dynamic Layer Tying for Parameter-Efficient TransformersTamir David Hay, Lior WolfICLR 2024 · 被引用 13 次
- Even Sparser Graph TransformersHamed Shirzad, Honghao Lin, Balaji Venkatachalam, Ameya Velingker 等NeurIPS 2024 · 被引用 18 次
