AutoGT: Automated Graph Transformer Architecture Search
Zizhao Zhang, Xin Wang, Chaoyu Guan, Ziwei Zhang, Haoyang Li, Wenwu Zhu
Abstract
Although Transformer architectures have been successfully applied to graph data with the advent of Graph Transformer, current design of Graph Transformer still heavily relies on human labor and expertise knowledge to decide proper neural architectures and suitable graph encoding strategies at each Transformer layer. In literature, there have been some works on automated design of Transformers focusing on non-graph data such as texts and images without considering graph encoding strategies, which fail to handle the non-euclidean graph data. In this paper, we study the problem of automated graph Transformer, for the first time. However, solving these problems poses the following challenges: i) how can we design a unified search space for graph Transformer, and ii) how to deal with the coupling relations between Transformer architectures and the graph encodings of each Transformer layer. To address these challenges, we propose Automated Graph Transformer (AutoGT), a neural architecture search framework that can automatically discover the optimal graph Transformer architectures by joint optimization of Transformer architecture and graph encoding strategies. Specifically, we first propose a unified graph Transformer formulation that can represent most of state-of-the-art graph Transformer architectures. Based upon the unified formulation, we further design the graph Transformer search space that includes both candidate architectures and various graph encodings. To handle the coupling relations, we propose a novel encoding-aware performance estimation strategy by gradually training and splitting the supernets according to the correlations between graph encodings and architectures. The proposed strategy can provide a more consistent and fine-grained performance prediction when evaluating the jointly optimized graph encodings and architectures. Extensive experiments and ablation studies show that our proposed AutoGT gains sufficient improvement over state-of-the-art hand-crafted baselines on all datasets, demonstrating its effectiveness and wide applicability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers9
- Adaptive Disentangled Transformer for Sequential RecommendationYipeng Zhang, Xin Wang, Hong Chen, Wenwu ZhuKDD 2023 · 32 citations
- Multi-task Graph Neural Architecture Search with Task-aware Collaboration and CurriculumYijian Qin, Xin Wang, Ziwei Zhang, Hong Chen et al.NeurIPS 2023 · 27 citations
- What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional EncodingHongkang Li, Meng Wang, Tengfei Ma, Sijia Liu et al.ICML 2024 · 23 citations
- Customized Subgraph Selection and Encoding for Drug-drug Interaction PredictionHaotong Du, Quanming Yao, Juzheng Zhang, Yang Liu et al.NeurIPS 2024 · 23 citations
- TorchGT: A Holistic System for Large-Scale Graph Transformer TrainingMeng Zhang, Jie Sun, Qinghao Hu, Peng Sun et al.SC 2024 · 7 citations
Related papers
- AutoGEL: An Automated Graph Neural Network with Explicit Link InformationZhili Wang, Shimin Di, Lei ChenNeurIPS 2021 · 46 citations
- Dynamic Heterogeneous Graph Attention Neural Architecture SearchZeyang Zhang, Ziwei Zhang, Xin Wang, Yijian Qin et al.AAAI 2023 · 44 citations
- AutoGSR: Neural Architecture Search for Graph-based Session RecommendationJingfan Chen, Guanghui Zhu, Haojun Hou, Chunfeng Yuan et al.SIGIR 2022 · 25 citations
- Graph External Attention Enhanced TransformerJianqing Liang, Min Chen, Jiye LiangICML 2024 · 11 citations
- Hypergraph Neural Architecture SearchWei Lin, Xu Peng, Zhengtao Yu, Taisong JinAAAI 2024 · 3 citations
