Viewing Transformers Through the Lens of Long Convolutions Layers
Itamar Zimerman, Lior Wolf
摘要
Despite their dominance in modern DL and, especially, NLP domains, transformer architectures exhibit sub-optimal performance on long-range tasks compared to recent layers that are specifically designed for this purpose. In this work, drawing inspiration from key attributes of longrange layers, such as state-space layers, linear RNN layers, and global convolution layers, we demonstrate that minimal modifications to the transformer architecture can significantly enhance performance on the Long Range Arena (LRA) benchmark, thus narrowing the gap with these specialized layers. We identify that two key principles for long-range tasks are (i) incorporating an inductive bias towards smoothness, and (ii) locality. As we show, integrating these ideas into the attention mechanism improves results with a negligible amount of additional computation and without any additional trainable parameters. Our theory and experiments also shed light on the reasons for the inferior performance of transformers on long-range tasks and identify critical properties that are essential for successfully capturing long-range dependencies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TensorLens: End-to-End Transformer Analysis via High-Order Attention TensorsIdo Andrew Atad, Itamar Zimerman, Shahar Katz, Lior WolfACL 2026 · 被引用 1 次
- Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention FormulationItamar Zimerman, Ameen Ali, Lior WolfICLR 2025
- Forgetting Transformer: Softmax Attention with a Forget GateZhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron C. CourvilleICLR 2025
它引用的顶会 Paper28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- FEDformer: Frequency Enhanced Decomposed Transformer for Long-term Series ForecastingTian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang 等ICML 2022 · 被引用 2,912 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- Never Train from Scratch: Fair Comparison of Long-Sequence Models Requires Data-Driven PriorsIdo Amos, Jonathan Berant, Ankit GuptaICLR 2024 · 被引用 39 次
- What Makes Convolutional Models Great on Long Sequence Modeling?Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen 等ICLR 2023 · 被引用 20 次
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for SequencesZhenhai Zhu, Radu SoricutACL 2021
- Reparameterized Multi-Resolution Convolutions for Long Sequence ModellingJake Cunningham, Giorgio Giannone, Mingtian Zhang, Marc Peter DeisenrothNeurIPS 2024 · 被引用 4 次
- Resurrecting Recurrent Neural Networks for Long SequencesAntonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando 等ICML 2023 · 被引用 474 次
