Relational Attention: Generalizing Transformers for Graph-Structured Tasks
Cameron Diao, Ricky Loynd
Abstract
Transformers flexibly operate over sets of real-valued vectors representing taskspecific entities and their attributes, where each vector might encode one wordpiece token and its position in a sequence, or some piece of information that carries no position at all. But as set processors, standard transformers are at a disadvantage in reasoning over more general graph-structured data where nodes represent entities and edges represent relations between entities. To address this shortcoming, we generalize transformer attention to consider and update edge vectors in each transformer layer. We evaluate this relational transformer on a diverse array of graph-structured tasks, including the large and challenging CLRS Algorithmic Reasoning Benchmark. There, it dramatically outperforms state-of-theart graph neural networks expressly designed to reason over graph-structured data. Our analysis demonstrates that these gains are attributable to relational attention's inherent ability to leverage the greater expressivity of graphs over sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2d410df2-83f4-46ae-91af-efdb27b3ac4fCited by top-tier papers25
- Graph Inductive Biases in Transformers without Message PassingLiheng Ma, Chen Lin, Derek Lim, Adriana Romero-Soriano et al.ICML 2023 · 185 citations
- Graph Neural Networks for Learning Equivariant Representations of Neural NetworksMiltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen et al.ICLR 2024 · 57 citations
- Transformers Meet Directed GraphsSimon Geisler, Yujia Li, Daniel J. Mankowitz, Ali Taylan Cemgil et al.ICML 2023 · 51 citations
- Graph Metanetworks for Processing Diverse Neural ArchitecturesDerek Lim, Haggai Maron, Marc T. Law, Jonathan Lorraine et al.ICLR 2024 · 47 citations
- Neural Algorithmic Reasoning with Causal RegularisationBeatrice Bevilacqua, Kyriacos Nikiforou, Borja Ibarz, Ioana Bica et al.ICML 2023 · 39 citations
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
Related papers
- Systematic Generalization with Edge TransformersLeon Bergen, Timothy J. O'Donnell, Dzmitry BahdanauNeurIPS 2021 · 62 citations
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio et al.ICLR 2024 · 21 citations
- Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in TransformersAwni Altabaa, Taylor Whittington Webb, Jonathan D. Cohen, John LaffertyICLR 2024 · 13 citations
- Disentangling and Integrating Relational and Sensory Information in Transformer ArchitecturesAwni Altabaa, John LaffertyICML 2025
- Generalizable Insights for Graph Transformers in Theory and PracticeTimo Stoll, Luis Müller, Christopher MorrisNeurIPS 2025 · 2 citations
