Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in Transformers
Awni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, John Lafferty
摘要
An extension of Transformers is proposed that enables explicit relational reasoning through a novel module called the Abstractor. At the core of the Abstractor is a variant of attention called relational cross-attention. The approach is motivated by an architectural inductive bias for relational learning that disentangles relational information from object-level features. This enables explicit relational reasoning, supporting abstraction and generalization from limited data. The Abstractor is first evaluated on simple discriminative relational tasks and compared to existing relational architectures. Next, the Abstractor is evaluated on purely relational sequence-to-sequence tasks, where dramatic improvements are seen in sample efficiency compared to standard Transformers. Finally, Abstractors are evaluated on a collection of tasks based on mathematical problem solving, where consistent improvements in performance and sample efficiency are observed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language ModelsYukang Yang, Declan Campbell, Kaixuan Huang, Mengdi Wang 等ICML 2025
- Mind the GAP: Glimpse-based Active Perception improves generalization and sample efficiency of visual reasoningOleh Kolner, Thomas Ortner, Stanislaw Wozniak, Angeliki PantaziICLR 2025
- Disentangling and Integrating Relational and Sensory Information in Transformer ArchitecturesAwni Altabaa, John LaffertyICML 2025
它引用的顶会 Paper6
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das 等NeurIPS 2023 · 被引用 444 次
- An Explicitly Relational Neural Network ArchitectureMurray Shanahan, Kyriacos Nikiforou, Antonia Creswell, Christos Kaplanis 等ICML 2020 · 被引用 72 次
- Emergent Symbols through Binding in External MemoryTaylor Whittington Webb, Ishan Sinha, Jonathan D. CohenICLR 2021 · 被引用 68 次
- Untangling tradeoffs between recurrence and self-attention in artificial neural networksGiancarlo Kerg, Bhargav Kanuparthi, Anirudh Goyal, Kyle Goyette 等NeurIPS 2020 · 被引用 16 次
相关 Paper
- Slot Abstractors: Toward Scalable Abstract Visual ReasoningShanka Subhra Mondal, Jonathan D. Cohen, Taylor Whittington WebbICML 2024 · 被引用 10 次
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio 等ICLR 2024 · 被引用 21 次
- Relational Attention: Generalizing Transformers for Graph-Structured TasksCameron Diao, Ricky LoyndICLR 2023 · 被引用 6 次
- Systematic Visual Reasoning through Object-Centric Relational AbstractionTaylor W. Webb, Shanka Subhra Mondal, Jonathan D. CohenNeurIPS 2023 · 被引用 35 次
- Enhancing Transformers for Generalizable First-Order Logical EntailmentTianshi Zheng, Jiazheng Wang, Zihao Wang, Jiaxin Bai 等ACL 2025 · 被引用 7 次
