DUALFormer: Dual Graph Transformer
Jiaming Zhuo, Yuwei Liu, Yintong Lu, Ziyi Ma, Kun Fu, Chuan Wang, Yuanfang Guo, Zhen Wang, Xiaochun Cao, Liang Yang
Abstract
Graph Transformers (GTs), adept at capturing the locality and globality of graphs, have shown promising potential in node classification tasks. Most state-of-the-art GTs succeed through integrating local Graph Neural Networks (GNNs) with their global Self-Attention (SA) modules to enhance structural awareness. Nonetheless, this architecture faces limitations arising from scalability challenges and the tradeoff between capturing local and global information. On the one hand, the quadratic complexity associated with the SA modules poses a significant challenge for many GTs, particularly when scaling them to large-scale graphs. Numerous GTs necessitated a compromise, relinquishing certain aspects of their expressivity to garner computational efficiency. On the other hand, GTs face challenges in maintaining detailed local structural information while capturing long-range dependencies. As a result, they typically require significant computational costs to balance the local and global expressivity. To address these limitations, this paper introduces a novel GT architecture, dubbed DUALFormer, featuring a dual-dimensional design of its GNN and SA modules. Leveraging approximation theory from Linearized Transformers and treating the query as the surrogate representation of node features, DUALFormer efficiently performs the computationally intensive global SA module on feature dimensions. Furthermore, by such a separation of local and global modules into dual dimensions, DUALFormer achieves a natural balance between local and global expressivity. In theory, DUALFormer can reduce intra-class variance, thereby enhancing the discriminability of node representations. Extensive experiments on eleven real-world datasets demonstrate its effectiveness and efficiency over existing state-of-the-art GTs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 054c37a2-7458-4e7c-8dfd-7db5ae160283Cited by top-tier papers15
- Rethinking Tokenized Graph Transformers for Node ClassificationJinsong Chen, Chenyang Li, Gaichao Li, John E. Hopcroft et al.NeurIPS 2025 · 8 citations
- A Closer Look at Graph Transformers: Cross-Aggregation and BeyondJiaming Zhuo, Ziyi Ma, Yintong Lu, Yuwei Liu et al.NeurIPS 2025 · 4 citations
- Graph Domain Adaptation via Homophily-Agnostic Reconstructing StructureRuiyi Fang, Shuo Wang, Ruizhi Pu, Qiuhao Zeng et al.AAAI 2026 · 1 citation
- Do We Really Need Message Passing in Brain Network Modeling?Liang Yang, Yuwei Liu, Jiaming Zhuo, Di Jin et al.ICML 2025
- On the Spectral Unreachability of Brain Graph LearningJiaming Zhuo, Shuai Zhai, Ziyi Ma, Kun Fu et al.ICML 2026
Builds on18
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu et al.NeurIPS 2022 · 1,216 citations
- Graph Random Neural Networks for Semi-Supervised Learning on GraphsWenzheng Feng, Jie Zhang, Yuxiao Dong, Yu Han et al.NeurIPS 2020 · 526 citations
- NodeFormer: A Scalable Graph Structure Learning Transformer for Node ClassificationQitian Wu, Wentao Zhao, Zenan Li, David P. Wipf et al.NeurIPS 2022 · 472 citations
- Representing Long-Range Context for Graph Neural Networks with Global AttentionZhanghao Wu, Paras Jain, Matthew A. Wright, Azalia Mirhoseini et al.NeurIPS 2021 · 450 citations
Related papers
- Primphormer: Efficient Graph Transformers with Primal RepresentationsMingzhen He, Ruikai Yang, Hanling Tian, Youmei Qiu et al.ICML 2025
- Polynormer: Polynomial-Expressive Graph Transformer in Linear TimeChenhui Deng, Zichao Yue, Zhiru ZhangICLR 2024 · 81 citations
- A Scalable and Effective Alternative to Graph TransformersKaan Sancak, Zhigang Hua, Jin Fang, Yan Xie et al.AAAI 2025 · 5 citations
- Distinguished In Uniform: Self-Attention Vs. Virtual NodesEran Rosenbluth, Jan Tönshoff, Martin Ritzert, Berke Kisin et al.ICLR 2024 · 20 citations
- HubGT: Fast Graph Transformer with Decoupled Hierarchy LabelingNingyi Liao, Zihao Yu, Siqiang Luo, Gao CongNeurIPS 2025 · 3 citations
