Less is More: on the Over-Globalizing Problem in Graph Transformers
Yujie Xing, Xiao Wang, Yibo Li, Hai Huang, Chuan Shi
Abstract
Graph Transformer, due to its global attention mechanism, has emerged as a new tool in dealing with graph-structured data. It is well recognized that the global attention mechanism considers a wider receptive field in a fully connected graph, leading many to believe that useful information can be extracted from all the nodes. In this paper, we challenge this belief: does the globalizing property always benefit Graph Transformers? We reveal the over-globalizing problem in Graph Transformer by presenting both empirical evidence and theoretical analysis, i.e., the current attention mechanism overly focuses on those distant nodes, while the near nodes, which actually contain most of the useful information, are relatively weakened. Then we propose a novel Bi-Level Global Graph Transformer with Collaborative Training (CoBFormer), including the intercluster and intra-cluster Transformers, to prevent the over-globalizing problem while keeping the ability to extract valuable information from distant nodes. Moreover, the collaborative training is proposed to improve the model's generalization ability with a theoretical guarantee. Extensive experiments on various graphs well validate the effectiveness of our proposed CoBFormer. The source code is available for reproducibility at: https://github.com/null-xyj/CoBFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b84338e8-100a-4288-9908-6e83c8591a2aCited by top-tier papers28
- Enhancing Graph Transformers with Hierarchical Distance Structural EncodingYuankai Luo, Hongkang Li, Lei Shi, Xiao-Ming WuNeurIPS 2024 · 26 citations
- FairGP: A Scalable and Fair Graph Transformer Using Graph PartitioningRenqiang Luo, Huafei Huang, Ivan Lee, Chengpei Xu et al.AAAI 2025 · 20 citations
- Deeper with Riemannian Geometry: Overcoming Oversmoothing and Oversquashing for Graph Foundation ModelsLi Sun, Zhenhao Huang, Ming Zhang, Philip S. YuNeurIPS 2025 · 10 citations
- Rethinking Tokenized Graph Transformers for Node ClassificationJinsong Chen, Chenyang Li, Gaichao Li, John E. Hopcroft et al.NeurIPS 2025 · 8 citations
- GaitCycFormer: Leveraging Gait Cycles and Transformers for Gait Emotion RecognitionQingyang Zeng, Lin ShangAAAI 2025 · 6 citations
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
Related papers
- Relieving the Over-Aggregating Effect in Graph TransformersJunshu Sun, Wanxing Chang, Chenxue Yang, Qingming Huang et al.NeurIPS 2025 · 3 citations
- Cooperative Graph Transformer with Structural Consensus for Multi-View LearningZhiyuan Lai, Jiacheng Li, Jiayuan Wang, Shiping WangAAAI 2026
- Tokenphormer: Structure-aware Multi-token Graph Transformer for Node ClassificationZijie Zhou, Zhaoqi Lu, Xuekai Wei, Rongqin Chen et al.AAAI 2025 · 5 citations
- HINormer: Representation Learning On Heterogeneous Information Networks with Graph TransformerQiheng Mao, Zemin Liu, Chenghao Liu, Jianling SunWWW 2023 · 106 citations
- Hierarchical Graph Transformer with Adaptive Node SamplingZaixi Zhang, Qi Liu, Qingyong Hu, Chee-Kong LeeNeurIPS 2022 · 145 citations
