Hierarchical Graph Transformer with Adaptive Node Sampling
Zaixi Zhang, Qi Liu, Qingyong Hu, Chee-Kong Lee
Abstract
The Transformer architecture has achieved remarkable success in a number of domains including natural language processing and computer vision. However, when it comes to graph-structured data, transformers have not achieved competitive performance, especially on large graphs. In this paper, we identify the main deficiencies of current graph transformers: (1) Existing node sampling strategies in Graph Transformers are agnostic to the graph characteristics and the training process. (2) Most sampling strategies only focus on local neighbors and neglect the long-range dependencies in the graph. We conduct experimental investigations on synthetic datasets to show that existing sampling strategies are sub-optimal. To tackle the aforementioned problems, we formulate the optimization strategies of node sampling in Graph Transformer as an adversary bandit problem, where the rewards are related to the attention weights and can vary in the training procedure. Meanwhile, we propose a hierarchical attention scheme with graph coarsening to capture the long-range interactions while reducing computational complexity. Finally, we conduct extensive experiments on real-world datasets to demonstrate the superiority of our method over existing graph transformers and popular GNNs. Our code is open-sourced at https://github.com/zaixizhang/ANS-GT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b642ec6-6d20-4087-8711-b354b9946cd3Cited by top-tier papers36
- A Generalization of ViT/MLP-Mixer to GraphsXiaoxin He, Bryan Hooi, Thomas Laurent, Adam Perold et al.ICML 2023 · 135 citations
- Forest-Based Graph Learning for Semi-Supervised Node ClassificationJin Li, Shenghao Gao, Kaichen Zhang, Xinlong Chen et al.ICLR 2026 · 132 citations
- Polynormer: Polynomial-Expressive Graph Transformer in Linear TimeChenhui Deng, Zichao Yue, Zhiru ZhangICLR 2024 · 81 citations
- Graph Mamba: Towards Learning on Graphs with State Space ModelsAli Behrouz, Farnoosh HashemiKDD 2024 · 63 citations
- VCR-Graphormer: A Mini-batch Graph Transformer via Virtual ConnectionsDongqi Fu, Zhigang Hua, Yan Xie, Jin Fang et al.ICLR 2024 · 47 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Beyond Homophily in Graph Neural Networks: Current Limitations and Effective DesignsJiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann et al.NeurIPS 2020 · 1,490 citations
Related papers
- Restricted Global-Aware Graph Filters Bridging GNNs and Transformer for Node ClassificationJingyuan Zhang, Xin Wang, Lei Yu, Zhirong Huang et al.NeurIPS 2025 · 3 citations
- A Scalable and Effective Alternative to Graph TransformersKaan Sancak, Zhigang Hua, Jin Fang, Yan Xie et al.AAAI 2025 · 5 citations
- Representing Long-Range Context for Graph Neural Networks with Global AttentionZhanghao Wu, Paras Jain, Matthew A. Wright, Azalia Mirhoseini et al.NeurIPS 2021 · 450 citations
- Less is More: on the Over-Globalizing Problem in Graph TransformersYujie Xing, Xiao Wang, Yibo Li, Hai Huang et al.ICML 2024
- NAGphormer: A Tokenized Graph Transformer for Node Classification in Large GraphsJinsong Chen, Kaiyuan Gao, Gaichao Li, Kun HeICLR 2023 · 22 citations
