Relieving the Over-Aggregating Effect in Graph Transformers
Junshu Sun, Wanxing Chang, Chenxue Yang, Qingming Huang, Shuhui Wang
Abstract
Graph attention has demonstrated superior performance in graph learning tasks. However, learning from global interactions can be challenging due to the large number of nodes. In this paper, we discover a new phenomenon termed over-aggregating. Over-aggregating arises when a large volume of messages is aggregated into a single node with less discrimination, leading to the dilution of the key messages and potential information loss. To address this, we propose Wideformer, a plug-and-play method for graph attention. Wideformer divides the aggregation of all nodes into parallel processes and guides the model to focus on specific subsets of these processes. The division can limit the input volume per aggregation, avoiding message dilution and reducing information loss. The guiding step sorts and weights the aggregation outputs, prioritizing the informative messages. Evaluations show that Wideformer can effectively mitigate over-aggregating. As a result, the backbone methods can focus on the informative messages, achieving superior performance compared to baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a17e0166-241d-42ce-9042-cc7d44582b8aCited by top-tier papers4
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept SetShufan Shen, Junshu Sun, Qingming Huang, Shuhui WangNeurIPS 2025 · 13 citations
- Edit Less, Achieve More: Dynamic Sparse Neuron Masking for Lifelong Knowledge Editing in LLMsJinzhe Liu, Junshu Sun, Shufan Shen, Chenxue Yang et al.NeurIPS 2025 · 8 citations
- Adaptive Recurrent Message Passing for Test Time Computing on GraphsJunshu Sun, Wanxing Chang, Qingming Huang, Shuhui WangICML 2026
- Enhancing LLMs for Graph Tasks via Graph-aware LoRA GenerationJunshu Sun, Wanxing Chang, Qingming Huang, Shuhui WangICML 2026
Builds on38
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
Related papers
- Less is More: on the Over-Globalizing Problem in Graph TransformersYujie Xing, Xiao Wang, Yibo Li, Hai Huang et al.ICML 2024
- A Closer Look at Graph Transformers: Cross-Aggregation and BeyondJiaming Zhuo, Ziyi Ma, Yintong Lu, Yuwei Liu et al.NeurIPS 2025 · 4 citations
- PolyFormer: Scalable Node-wise Filters via Polynomial Graph TransformerJiahong Ma, Mingguo He, Zhewei WeiKDD 2024 · 6 citations
- Towards Deeper Graph Neural NetworksMeng Liu, Hongyang Gao, Shuiwang JiKDD 2020 · 496 citations
- Even Sparser Graph TransformersHamed Shirzad, Honghao Lin, Balaji Venkatachalam, Ameya Velingker et al.NeurIPS 2024 · 18 citations
