Cluster-wise Graph Transformer with Dual-granularity Kernelized Attention
Siyuan Huang, Yunchong Song, Jiayue Zhou, Zhouhan Lin
Abstract
In the realm of graph learning, there is a category of methods that conceptualize graphs as hierarchical structures, utilizing node clustering to capture broader structural information. While generally effective, these methods often rely on a fixed graph coarsening routine, leading to overly homogeneous cluster representations and loss of node-level information. In this paper, we envision the graph as a network of interconnected node sets without compressing each cluster into a single embedding. To enable effective information transfer among these node sets, we propose the Node-to-Cluster Attention (N2C-Attn) mechanism. N2C-Attn incorporates techniques from Multiple Kernel Learning into the kernelized attention framework, effectively capturing information at both node and cluster levels. We then devise an efficient form for N2C-Attn using the cluster-wise message-passing framework, achieving linear time complexity. We further analyze how N2C-Attn combines bi-level feature maps of queries and keys, demonstrating its capability to merge dual-granularity information. The resulting architecture, Cluster-wise Graph Transformer (Cluster-GT), which uses node clusters as tokens and employs our proposed N2C-Attn module, shows superior performance on various graph-level tasks. Code is available at https://github.com/LUMIA-Group/Cluster-wise-Graph-Transformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd293faf-4e6d-42c1-bc50-63d6c94403baCited by top-tier papers4
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask FrameworkYujie Xing, Xiao Wang, Bin Wu, Hai Huang et al.NeurIPS 2025 · 2 citations
- Can Classic GNNs Be Strong Baselines for Graph-level Tasks? Simple Architectures Meet ExcellenceYuankai Luo, Lei Shi, Xiao-Ming WuICML 2025
- Dual-Kernel Graph Community Contrastive LearningXiang Chen, Kun Yue, Wenjie Liu, Zhenyu Zhang et al.AAAI 2026
- From atom to space: A region-based readout function for spatial properties of materialsJiawen Zou, Weimin Tan, Zhongyao Wang, Hao Qi et al.ICLR 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
Related papers
- Deep Multi-modal Graph Clustering via Graph Transformer NetworkQianqian Wang, Haiming Xu, Zihao Zhang, Wei Feng et al.AAAI 2025 · 4 citations
- Anchor-Driven Nyström for Deep Graph-Level ClusteringJiaxin Wang, Wenxuan Tu, Lingren Wang, Jieren Cheng et al.AAAI 2026
- A Unified Graph Clustering NetworkRenda Han, Xiaobao Wang, Longbiao Wang, Wenxin Zhang et al.WWW 2026
- ClusterGNN: Cluster-based Coarse-to-Fine Graph Neural Network for Efficient Feature MatchingYan Shi, Junxiong Cai, Yoli Shavit, Tai-Jiang Mu et al.CVPR 2022 · 91 citations
- Attention-driven Graph Clustering NetworkZhihao Peng, Hui Liu, Yuheng Jia, Junhui HouACM MM 2021 · 135 citations
