Associative Transformer
Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin, Ryota Kanai
Abstract
Emerging from the pairwise attention in conventional Transformers, there is a growing interest in sparse attention mechanisms that align more closely with localized, contextual learning in the biological brain. Existing studies such as the Coordination method employ iterative cross-attention mechanisms with a bottleneck to enable the sparse association of inputs. However, these methods are parameter inefficient and fail in more complex relational reasoning tasks. To this end, we propose Associative Transformer (AiT) to enhance the association among sparsely attended input tokens, improving parameter efficiency and performance in various vision tasks such as classification and relational reasoning. AiT leverages a learnable explicit memory comprising specialized priors that guide bottleneck attentions to facilitate the extraction of diverse localized tokens. Moreover, AiT employs an associative memory-based token reconstruction using a Hopfield energy function. The extensive empirical experiments demonstrate that AiT requires significantly fewer parameters and attention layers outperforming a broad range of sparse Transformer models. Additionally, AiT outperforms the SOTA sparse Transformer models including the Coordination method on the Sort-of-CLEVR dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch et al.ICLR 2022 · 797 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
- Taming Sparsely Activated Transformer with Stochastic ExpertsSimiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim et al.ICLR 2022 · 144 citations
Related papers
- In-Context Compositional Learning vis Sparse Coding TransformerWei Chen, Jingxi Yu, Zichen Miao, Qiang QiuNeurIPS 2025
- BiFormer: Vision Transformer with Bi-Level Routing AttentionLei Zhu, Xinjiang Wang, Zhanghan Ke, Wayne Zhang et al.CVPR 2023
- Disentangling and Integrating Relational and Sensory Information in Transformer ArchitecturesAwni Altabaa, John LaffertyICML 2025
- Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in TransformersAwni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, John LaffertyICLR 2024 · 13 citations
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio et al.ICLR 2024 · 21 citations
