FiX: Introducing Fine-grained Forget Gate into Softmax Attention
Runzhong Li, Renjie Liu, Qing Li, Bo Tang
Abstract
Causal softmax attention is the algorithmic foundation of modern large language models. Inspired by linear attention, recent work has sought to enhance it by incorporating forget gates. However, these efforts, such as FoX, have been limited to coarse, scalar gates. While fine-grained, element-wise gates are shown to be more effective than scalar ones in linear attention, their direct integration into softmax attention is non-trivial due to algebraic constraints. In this work, we introduce Fine-grained Forgetting Transformer ( FiX ), a novel architecture that successfully enables element-wise forget gates in softmax attention. Our core insight is that the softmax denominator becomes mathematically redundant under a subsequent RMSNorm layer, allowing us to reformulate the forgetting mechanism as a direct element-wise multiplication on the value vectors. This formulation makes FiX the first positional encoding applied to value-output (VO) pairs, designed to be complementary to existing query-key (QK) encodings like RoPE. We systematically address implementation challenges including numerical precision, computational efficiency, and inference memory consumption. Extensive experiments show FiX achieves lower training loss and superior performance on both short-text common sense benchmarks and long-context tasks, opening a new path for building more powerful transformers. Our code is available at https://github.com/AlayaDB-AI/FiX.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fdc9702-76da-4293-a923-0c969fee1939Builds on21
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Selective Rotary Position EmbeddingSajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter et al.ICLR 2026 · 11 citations
- Forgetting Transformer: Softmax Attention with a Forget GateZhixuan Lin, Evgenii Nikishin, Xu Owen He, Aaron C. CourvilleICLR 2025
- Group Representational Position EncodingYifan Zhang, Zixiang Chen, Yifeng Liu, Zhen Qin et al.ICLR 2026 · 5 citations
- Frayed RoPE and Long Inputs: A Geometric PerspectiveDavis Wertheimer, Aozhong Zhang, Derrick Liu, Penghang Yin et al.ICLR 2026 · 3 citations
- PaTH Attention: Position Encoding via Accumulating Householder TransformationsSonglin Yang, Yikang Shen, Kaiyue Wen, Shawn Tan et al.NeurIPS 2025 · 36 citations
