Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in Transformers
Awni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, John Lafferty
Abstract
An extension of Transformers is proposed that enables explicit relational reasoning through a novel module called the Abstractor. At the core of the Abstractor is a variant of attention called relational cross-attention. The approach is motivated by an architectural inductive bias for relational learning that disentangles relational information from object-level features. This enables explicit relational reasoning, supporting abstraction and generalization from limited data. The Abstractor is first evaluated on simple discriminative relational tasks and compared to existing relational architectures. Next, the Abstractor is evaluated on purely relational sequence-to-sequence tasks, where dramatic improvements are seen in sample efficiency compared to standard Transformers. Finally, Abstractors are evaluated on a collection of tasks based on mathematical problem solving, where consistent improvements in performance and sample efficiency are observed.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06e777d0-99e8-4488-99fc-be5d1517e502Cited by top-tier papers3
- Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language ModelsYukang Yang, Declan Campbell, Kaixuan Huang, Mengdi Wang et al.ICML 2025
- Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoningOleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki PantaziICLR 2025
- Disentangling and Integrating Relational and Sensory Information in Transformer ArchitecturesAwni Altabaa, John LaffertyICML 2025
Builds on6
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
- An Explicitly Relational Neural Network ArchitectureMurray Shanahan, Kyriacos Nikiforou, Antonia Creswell, Christos Kaplanis et al.ICML 2020 · 72 citations
- Emergent Symbols through Binding in External MemoryTaylor Whittington Webb, Ishan Sinha, Jonathan D. CohenICLR 2021 · 68 citations
- Untangling tradeoffs between recurrence and self-attention in artificial neural networksGiancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal, Kyle Goyette et al.NeurIPS 2020 · 16 citations
Related papers
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 10 citations
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio et al.ICLR 2024 · 21 citations
- Relational Attention: Generalizing Transformers for Graph-Structured TasksCameron Diao, Ricky LoyndICLR 2023 · 6 citations
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 35 citations
- Enhancing Transformers for Generalizable First-Order Logical EntailmentTianshi Zheng, Jiazheng Wang, Zihao Wang, Jiaxin Bai et al.ACL 2025 · 7 citations
