Compositional Attention: Disentangling Search and Retrieval
Sarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio, Guillaume Lajoie
摘要
Multi-head, key-value attention is the backbone of the widely successful Transformer model and its variants. This attention mechanism uses multiple parallel key-value attention blocks (called heads), each performing two fundamental computations: (1) search - selection of a relevant entity from a set via query-key interactions, and (2) retrieval - extraction of relevant features from the selected entity via a value matrix. Importantly, standard attention heads learn a rigid mapping between search and retrieval. In this work, we first highlight how this static nature of the pairing can potentially: (a) lead to learning of redundant parameters in certain tasks, and (b) hinder generalization. To alleviate this problem, we propose a novel attention mechanism, called Compositional Attention, that replaces the standard head structure. The proposed mechanism disentangles search and retrieval and composes them in a dynamic, flexible and context-dependent manner through an additional soft competition stage between the query-key combination and value pairing. Through a series of numerical experiments, we show that it outperforms standard multi-head attention on a variety of tasks, including some out-of-distribution settings. Through our qualitative analysis, we demonstrate that Compositional Attention leads to dynamic specialization based on the type of retrieval needed. Our proposed mechanism generalizes multi-head attention, allows independent scaling of search and retrieval, and can easily be implemented in lieu of standard attention heads in any network architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Is a Modular Architecture Enough?Sarthak Mittal, Yoshua Bengio, Guillaume LajoieNeurIPS 2022 · 被引用 62 次
- Improving Transformers with Probabilistic Attention KeysTam Minh Nguyen, Tan Minh Nguyen, Dung D. Le, Duy Khuong Nguyen 等ICML 2022 · 被引用 38 次
- Layer-Wise Representation Fusion for Compositional GeneralizationYafang Zheng, Lei Lin, Shuangtao Li, Yuxuan Yuan 等AAAI 2024 · 被引用 4 次
- Attention-based Iterative Decomposition for Tensor Product RepresentationTaewon Park, Inchul Choi, Minho LeeICLR 2024 · 被引用 1 次
- Discrete Dictionary-based Decomposition Layer for Structured Representation LearningTaewon Park, Hyun-Chul Kim, Minho LeeNeurIPS 2024 · 被引用 1 次
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 等NeurIPS 2020 · 被引用 1,275 次
- A Meta-Transfer Objective for Learning to Disentangle Causal MechanismsYoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke 等ICLR 2020 · 被引用 371 次
- Rethinking Attention with PerformersKrzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song 等ICLR 2021 · 被引用 122 次
相关 Paper
- Attention as a HypernetworkSimon Schug, Seijin Kobayashi, Yassir Akram, João Sacramento 等ICLR 2025
- Improving Transformers with Dynamically Composable Multi-Head AttentionDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2024 · 被引用 8 次
- In-Context Compositional Learning vis Sparse Coding TransformerWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuNeurIPS 2025
- Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language ModelingXiang Hu, Zhihao Teng, Jun Zhao, Wei Wu 等ICML 2025
- Retrieval-Aware Distillation for Transformer-SSM HybridsAviv Bick, Eric Xing, Albert GuICML 2026 · 被引用 4 次
