What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
Hongkang Li, Meng Wang, Tengfei Ma, Sijia Liu, Zaixi Zhang, Pin-Yu Chen
Abstract
Graph Transformers, which incorporate self-attention and positional encoding, have recently emerged as a powerful architecture for various graph learning tasks. Despite their impressive performance, the complex non-convex interactions across layers and the recursive graph structure have made it challenging to establish a theoretical foundation for learning and generalization. This study introduces the first theoretical investigation of a shallow Graph Transformer for semi-supervised node classification, comprising a self-attention layer with relative positional encoding and a two-layer perceptron. Focusing on a graph data model with discriminative nodes that determine node labels and non-discriminative nodes that are class-irrelevant, we characterize the sample complexity required to achieve a desirable generalization error by training with stochastic gradient descent (SGD). This paper provides the quantitative characterization of the sample complexity and number of iterations for convergence dependent on the fraction of discriminative nodes, the dominant patterns, and the initial model errors. Furthermore, we demonstrate that self-attention and positional encoding enhance generalization by making the attention map sparse and promoting the core neighborhood during training, which explains the superior feature representation of Graph Transformers. Our theoretical results are supported by empirical experiments on synthetic and real-world benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6c9fbe69-0515-4f30-a3d6-c094cc6f49d2Cited by top-tier papers11
- Enhancing Graph Transformers with Hierarchical Distance Structural EncodingYuankai Luo, Hongkang Li, Lei Shi, Xiao-Ming WuNeurIPS 2024 · 26 citations
- How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?Wei Huang, Andi Han, Yujin Song, Yilan Chen et al.NeurIPS 2025 · 4 citations
- A Theoretical Analysis of Mamba’s Training Dynamics: Filtering Relevant Features for Generalization in State Space ModelsMugunthan Shandirasegaran, Hongkang Li, Songyang Zhang, Meng Wang et al.ICLR 2026 · 3 citations
- Generalizable Insights for Graph Transformers in Theory and PracticeTimo Stoll, Luis Müller, Christopher MorrisNeurIPS 2025 · 2 citations
- Invariant Graph Transformer for Out-of-Distribution GeneralizationTianyin Liao, Ziwei Zhang, Yufei Sun, Chunyu Hu et al.KDD 2026 · 2 citations
Builds on42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
Related papers
- Generalization Guarantee of Training Graph Convolutional Networks with Graph Topology SamplingHongkang Li, Meng Wang, Sijia Liu, Pin-Yu Chen et al.ICML 2022 · 34 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token AttentionJingbo Zhou, Jun Xia, Siyuan Li, Yunfan Liu et al.WWW 2026
- Are More Layers Beneficial to Graph Transformers?Haiteng Zhao, Shuming Ma, Dongdong Zhang, Zhi-Hong Deng et al.ICLR 2023 · 5 citations
- Graph Positional and Structural EncoderSemih Cantürk, Renming Liu, Olivier Lapointe-Gagné, Vincent Létourneau et al.ICML 2024 · 33 citations
