Head-in-Head in Linear Attention
Shijie Mei, Man Yao, Jiabo Tong, Bo XU, Guoqi Li
摘要
The state-transition (decay) matrix governs how fixed-size memory is updated and used, making it a core design in linear attention models. Prior work exploits rank-1 approximations to reduce the cost of constructing decay matrices, but this low-rank constraint also limits the expressive capacity. We therefore formulate decay-matrix design as an open optimization problem: maximizing expressiveness while introducing minimal additional cost. Inspired by the multi-head mechanism, we propose Head-in-Head, which introduces an additional mask matrix to structure memory partitioning and interactions within a single linear-attention head. This simple, generic, and efficient design: 1) enables a rank- approximation of the decay matrix with only a few extra parameters and 2) strengthens intra-head information interaction. We further develop mask normalization and a chunk-wise parallelization scheme to support efficient parallel training. Extensive experiments on synthetic benchmarks and language modeling tasks, together with visual analyses, show that Head-in-Head consistently improves baseline performance by enriching information diversity and strengthening intra-head interactions. Code available at: https://github.com/msj-19/Head-in-Head-Linear-Attention
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The emergence of sparse attention: impact of data distribution and benefits of repetitionNicolas Zucchet, Francesco D'Angelo, Andrew Kyle Lampinen, Stephanie ChanNeurIPS 2025 · 被引用 28 次
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural NetworksDaniel Kunin, Giovanni Luca Marchetti, Feng Chen, Dhruva Karkada 等NeurIPS 2025 · 被引用 15 次
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in TransformersYixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao 等NeurIPS 2025 · 被引用 10 次
- Sequential Group Composition: A Window into the Mechanics of Deep LearningGiovanni Luca Marchetti, Daniel Kunin, Adele Myers, Francisco Acosta 等ICML 2026 · 被引用 8 次
- What Happens During the Loss Plateau? Understanding Abrupt Learning in TransformersPulkit Gopalani, Wei HuNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper29
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Householder-Diagonalized Linear Attention (HDLA): Utilizing Enhanced Decay Mechanism for Efficient Sequence ModelingJiefu Zhang, Zhen Qin, Jiabo Tong, Shijie Mei 等ICLR 2026
- DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionYilong Chen, Linhao Zhang, Junyuan Shang, Zhenyu Zhang 等NeurIPS 2024 · 被引用 12 次
- MVA: Linear Attention with High-order Query-Keys Integration and Multi-level Vocabulary DecompositionNing Wang, Zekun Li, Tongxin Bai, Man Yao 等ICML 2025
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 被引用 8 次
- Scaling Linear Attention Capacity with Sparse State ExpansionYuqi Pan, Yongqi An, Zheng Li, Yuhong Chou 等ICLR 2026 · 被引用 3 次
