Attention as a Hypernetwork
Simon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento, Razvan Pascanu
Abstract
Transformers can under some circumstances generalize to novel problem instances whose constituent parts might have been encountered during training, but whose compositions have not. What mechanisms underlie this ability for compositional generalization? By reformulating multi-head attention as a hypernetwork, we reveal that a composable, low-dimensional latent code specifies key-query specific operations. We find empirically that this latent code is predictive of the subtasks the network performs on unseen task compositions, revealing that latent codes acquired during training are reused to solve unseen problem instances. To further examine the hypothesis that the intrinsic hypernetwork of multi-head attention supports compositional generalization, we ablate whether making the hypernetwork-generated linear value network nonlinear strengthens compositionality. We find that this modification improves compositional generalization on abstract reasoning tasks. In particular, we introduce a symbolic version of the Raven's Progressive Matrices human intelligence test, which gives us precise control over the problem compositions encountered during training and evaluation. We demonstrate on this task how scaling model size and data enables compositional generalization in transformers and gives rise to a functionally structured latent space. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4bbcb04a-f2ed-47b5-96db-2e1bac7a86f2Cited by top-tier papers8
- Scaling can lead to compositional generalizationFlorian Redhardt, Yassir Akram, Simon SchugNeurIPS 2025 · 11 citations
- Hyper-GoalNet: Goal-Conditioned Manipulation Policy Learning with HyperNetworksPei Zhou, Wanting Yao, Qian Luo, Xunzhe Zhou et al.NeurIPS 2025 · 4 citations
- Emergent Analogical Reasoning in TransformersGouki Minegishi, Jingyuan Feng, Hiroki Furuta, Takeshi Kojima et al.ICML 2026 · 4 citations
- LoRAGen: Structure-Aware Weight Space Learning for LoRA GenerationHao Huang, Jingtao Ding, Mengqi Liao, Xin Wang et al.ICLR 2026
- Text-to-LoRA: Instant Transformer AdaptionRujikorn Charakorn, Edoardo Cetin, Yujin Tang, Robert Tjarko LangeICML 2025
Builds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller et al.ICLR 2024 · 229 citations
- Exphormer: Sparse Transformers for GraphsHamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J. Sutherland et al.ICML 2023 · 219 citations
Related papers
- In-Context Compositional Learning vis Sparse Coding TransformerWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuNeurIPS 2025
- Compositional Attention: Disentangling Search and RetrievalSarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio et al.ICLR 2022 · 20 citations
- Dynamic Inference with Neural InterpretersNasim Rahaman, Muhammad Waleed Gondal, Shruti Joshi, Peter V. Gehler et al.NeurIPS 2021 · 35 citations
- Compositional Generalization through Gradient Search in Nonparametric Latent SpaceHaruki Shirakami, James HendersonICLR 2026
- Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?Yutong Yin, Zhaoran WangICLR 2025
