Attention as a Hypernetwork
Simon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento, Razvan Pascanu
摘要
Transformers can under some circumstances generalize to novel problem instances whose constituent parts might have been encountered during training, but whose compositions have not. What mechanisms underlie this ability for compositional generalization? By reformulating multi-head attention as a hypernetwork, we reveal that a composable, low-dimensional latent code specifies key-query specific operations. We find empirically that this latent code is predictive of the subtasks the network performs on unseen task compositions, revealing that latent codes acquired during training are reused to solve unseen problem instances. To further examine the hypothesis that the intrinsic hypernetwork of multi-head attention supports compositional generalization, we ablate whether making the hypernetwork-generated linear value network nonlinear strengthens compositionality. We find that this modification improves compositional generalization on abstract reasoning tasks. In particular, we introduce a symbolic version of the Raven's Progressive Matrices human intelligence test, which gives us precise control over the problem compositions encountered during training and evaluation. We demonstrate on this task how scaling model size and data enables compositional generalization in transformers and gives rise to a functionally structured latent space. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 被引用 11 次
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworksPei Zhou, Wanting Yao, Qian Luo, Xunzhe Zhou 等NeurIPS 2025 · 被引用 4 次
- Emergent Analogical Reasoning in TransformersGouki Minegishi, Jingyuan Feng, Hiroki Furuta, Takeshi Kojima 等ICML 2026 · 被引用 4 次
- LoRAGen: Structure-Aware Weight Space Learning for LoRA GenerationHao Huang, Jingtao Ding, Mengqi Liao, Xin Wang 等ICLR 2026
- Text-to-LoRA: Instant Transformer AdaptionRujikorn Charakorn, Edoardo Cetin, Yujin Tang, Robert Tjarko LangeICML 2025
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li 等NeurIPS 2023 · 被引用 728 次
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 被引用 394 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
- Exphormer: Sparse Transformers for GraphsHamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J. Sutherland 等ICML 2023 · 被引用 219 次
相关 Paper
- In-Context Compositional Learning vis Sparse Coding TransformerWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuNeurIPS 2025
- Compositional Attention: Disentangling Search and RetrievalSarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio 等ICLR 2022 · 被引用 20 次
- Dynamic Inference with Neural InterpretersNasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler 等NeurIPS 2021 · 被引用 35 次
- Compositional Generalization through Gradient Search in Nonparametric Latent SpaceHaruki Shirakami, James HendersonICLR 2026
- Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?Yutong Yin, Zhaoran WangICLR 2025
