Selective Attention Improves Transformer
Yaniv Leviathan, Matan Kalman, Yossi Matias
Abstract
Unneeded elements in the attention's context degrade performance. We introduce Selective Attention, a simple parameter-free change to the standard attention mechanism which reduces attention to unneeded elements. Selective attention consistently improves language modeling and downstream task performance in a variety of model sizes and context lengths. For example, transformers trained with the language modeling objective on C4 with selective attention perform language modeling equivalently to standard transformers with ∼2X more heads and parameters in their attention modules. Selective attention also allows decreasing the size of the attention's context buffer, leading to meaningful reductions in the memory and compute requirements during inference. For example, transformers trained on C4 with context sizes of 512, 1,024, and 2,048 need 16X, 25X, and 47X less memory for their attention module, respectively, when equipped with selective attention, as those without selective attention, with the same validation perplexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Do Language Models Use Their Depth Efficiently?Róbert Csordás, Christopher D. Manning, Christopher PottsNeurIPS 2025 · 61 citations
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin et al.NeurIPS 2025 · 51 citations
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan et al.NeurIPS 2025 · 36 citations
- VORTA: Efficient Video Diffusion via Routing Sparse AttentionWenhao Sun, Rong-Cheng Tu, Yifu Ding, Jingyi Liao et al.NeurIPS 2025 · 25 citations
- SPARTAN: A Sparse Transformer World Model Attending to What MattersAnson Lei, Bernhard Schölkopf, Ingmar PosnerNeurIPS 2025 · 12 citations
Builds on19
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Dynamic Context Pruning for Efficient and Interpretable Autoregressive TransformersSotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci et al.NeurIPS 2023 · 95 citations
- Selective Attention: Enhancing Transformer through Principled Context ControlXuechen Zhang, Xiangyu Chang, Mingchen Li, Amit K. Roy-Chowdhury et al.NeurIPS 2024 · 32 citations
- EL-Attention: Memory Efficient Lossless Attention for GenerationYu Yan, Jiusheng Chen, Weizhen Qi, Nikhil Bhendawade et al.ICML 2021 · 9 citations
- Layer-Condensed KV Cache for Efficient Inference of Large Language ModelsHaoyi Wu, Kewei TuACL 2024
- Compressing Context to Enhance Inference Efficiency of Large Language ModelsYucheng Li, Bo Dong, Frank Guerin, Chenghua LinEMNLP 2023 · 54 citations
