Head-in-Head in Linear Attention
Shijie Mei, Man Yao, Jiabo Tong, Bo XU, Guoqi Li
Abstract
The state-transition (decay) matrix governs how fixed-size memory is updated and used, making it a core design in linear attention models. Prior work exploits rank-1 approximations to reduce the cost of constructing decay matrices, but this low-rank constraint also limits the expressive capacity. We therefore formulate decay-matrix design as an open optimization problem: maximizing expressiveness while introducing minimal additional cost. Inspired by the multi-head mechanism, we propose Head-in-Head, which introduces an additional mask matrix to structure memory partitioning and interactions within a single linear-attention head. This simple, generic, and efficient design: 1) enables a rank- approximation of the decay matrix with only a few extra parameters and 2) strengthens intra-head information interaction. We further develop mask normalization and a chunk-wise parallelization scheme to support efficient parallel training. Extensive experiments on synthetic benchmarks and language modeling tasks, together with visual analyses, show that Head-in-Head consistently improves baseline performance by enriching information diversity and strengthening intra-head interactions. Code available at: https://github.com/msj-19/Head-in-Head-Linear-Attention
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18e27208-0789-499f-b0ce-ab2c33e53daeCited by top-tier papers7
- The emergence of sparse attention: impact of data distribution and benefits of repetitionNicolas Zucchet, Francesco D'Angelo, Andrew Kyle Lampinen, Stephanie ChanNeurIPS 2025 · 28 citations
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada et al.NeurIPS 2025 · 15 citations
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in TransformersYixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao et al.NeurIPS 2025 · 10 citations
- Sequential Group Composition: A Window into the Mechanics of Deep LearningGiovanni Luca Marchetti, Daniel Kunin, Adele Myers, Francisco Acosta et al.ICML 2026 · 8 citations
- What Happens During the Loss Plateau? Understanding Abrupt Learning in TransformersPulkit Gopalani, Wei HuNeurIPS 2025 · 6 citations
Builds on29
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Householder-Diagonalized Linear Attention (HDLA): Utilizing Enhanced Decay Mechanism for Efficient Sequence ModelingJiefu Zhang, Zhen Qin, Jiabo Tong, Shijie Mei et al.ICLR 2026
- DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionYilong Chen, Linhao Zhang, Junyuan Shang, Zhenyu Zhang et al.NeurIPS 2024 · 12 citations
- MVA: Linear Attention with High-order Query-Keys Integration and Multi-level Vocabulary DecompositionNing Wang, Zekun Li, Tongxin Bai, Man Yao et al.ICML 2025
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 8 citations
- Scaling Linear Attention Capacity with Sparse State ExpansionYuqi Pan, Yongqi An, Zheng Li, Yuhong Chou et al.ICLR 2026 · 3 citations
