Linear Log-Normal Attention with Unbiased Concentration
Yury Nahshan, Joseph Kampeas, Emir Haleva
摘要
Transformer models have achieved remarkable results in a wide range of applications. However, their scalability is hampered by the quadratic time and memory complexity of the self-attention mechanism concerning the sequence length. This limitation poses a substantial obstacle when dealing with long documents or high-resolution images. In this work, we study the self-attention mechanism by analyzing the distribution of the attention matrix and its concentration ability. Furthermore, we propose instruments to measure these quantities and introduce a novel self-attention mechanism, Linear Log-Normal Attention, designed to emulate the distribution and concentration behavior of the original self-attention. Our experimental results on popular natural language benchmarks reveal that our proposed Linear Log-Normal Attention outperforms other linearized attention alternatives, offering a promising avenue for enhancing the scalability of transformer models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Parallelizing Linear Transformers with the Delta Rule over Sequence LengthSonglin Yang, Bailin Wang, Yu Zhang, Yikang Shen 等NeurIPS 2024 · 被引用 412 次
- Gated Linear Attention Transformers with Hardware-Efficient TrainingSonglin Yang, Bailin Wang, Yikang Shen, Rameswar Panda 等ICML 2024 · 被引用 390 次
- Length-Induced Embedding Collapse in PLM-based ModelsYuqi Zhou, Sunhao Dai, Zhanshuo Cao, Xiao Zhang 等ACL 2025 · 被引用 8 次
- NormDirection: Restoring the Missing Query Norm in Vision Linear AttentionWeikang Meng, Yadan Luo, Liangyu Huo, Yingjian Li 等ICML 2026 · 被引用 2 次
- Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature ScalingSam Hilton-Jones, Timothy Norman, Zhanxing ZhuICML 2026
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- The Devil in Linear TransformerZhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li 等EMNLP 2022 · 被引用 24 次
- Long-Short Transformer: Efficient Transformers for Language and VisionChen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi 等NeurIPS 2021 · 被引用 180 次
- Luna: Linear Unified Nested AttentionXuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou 等NeurIPS 2021 · 被引用 145 次
- Log-Linear AttentionHan Guo, Songlin Yang, Tarushii Goel, Eric P. Xing 等ICLR 2026 · 被引用 41 次
- Bridging the Divide: Reconsidering Softmax and Linear AttentionDongchen Han, Yifan Pu, Zhuofan Xia, Yizeng Han 等NeurIPS 2024 · 被引用 69 次
