Are More Layers Beneficial to Graph Transformers?
Haiteng Zhao, Shuming Ma, Dongdong Zhang, Zhi-Hong Deng, Furu Wei
Abstract
Despite that going deep has proven successful in many neural architectures, the existing graph transformers are relatively shallow. In this work, we explore whether more layers are beneficial to graph transformers, and find that current graph transformers suffer from the bottleneck of improving performance by increasing depth. Our further analysis reveals the reason is that deep graph transformers are limited by the vanishing capacity of global attention, restricting the graph transformer from focusing on the critical substructure and obtaining expressive features. To this end, we propose a novel graph transformer model named DeepGraph that explicitly employs substructure tokens in the encoded representation, and applies local attention on related nodes to obtain substructure based attention encoding. Our model enhances the ability of the global attention to focus on substructures and promotes the expressiveness of the representations, addressing the limitation of self-attention as the graph transformer deepens. Experiments show that our method unblocks the depth limitation of graph transformers and results in state-of-the-art performance across various graph benchmarks with deeper models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 531073b4-efca-4d75-8f6c-d17ff1e09090Cited by top-tier papers9
- GIMLET: A Unified Graph-Text Model for Instruction-Based Molecule Zero-Shot LearningHaiteng Zhao, Shengchao Liu, Chang Ma, Hannan Xu et al.NeurIPS 2023 · 97 citations
- Polynormer: Polynomial-Expressive Graph Transformer in Linear TimeChenhui Deng, Zichao Yue, Zhiru ZhangICLR 2024 · 81 citations
- Towards Deep Attention in Graph Neural Networks: Problems and RemediesSoo Yong Lee, Fanchen Bu, Jaemin Yoo, Kijung ShinICML 2023 · 44 citations
- Transformers Get Stable: An End-to-End Signal Propagation Theory for Language ModelsAkhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung et al.ICML 2024 · 16 citations
- Understanding Catastrophic Forgetting In LoRA via Mean-Field Attention DynamicsHugo Koubbi, Louis Hernandez, Matthieu BoussardICML 2026 · 7 citations
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
Related papers
- Structure-Aware Transformer for Graph Representation LearningDexiong Chen, Leslie O'Bray, Karsten M. BorgwardtICML 2022 · 349 citations
- Tokenphormer: Structure-aware Multi-token Graph Transformer for Node ClassificationZijie Zhou, Zhaoqi Lu, Xuekai Wei, Rongqin Chen et al.AAAI 2025 · 5 citations
- Tailoring Self-Attention for Graph via Rooted SubtreesSiyuan Huang, Yunchong Song, Jiayue Zhou, Zhouhan LinNeurIPS 2023 · 12 citations
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 55 citations
- Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node ClassificationJingyuan Zhang, Xin Wang, Lei Yu, Zhirong Huang et al.NeurIPS 2025 · 3 citations
