Generalizing GNNs with Tokenized Mixture of Experts
Xiaoguang Guo, Zehong Wang, Jiazheng Li, Shawn Spitzel, Qi Yang, Kaize Ding, Jundong Li, Chuxu Zhang
Abstract
Deployed graph neural networks (GNNs) operate as frozen snapshots, yet must simultaneously fit clean data, generalize under distribution shifts, and remain stable against input perturbations---three goals that are difficult to satisfy at once with a single fixed model. We first show theoretically that a single fixed inference rule can create a stability--generalization tradeoff: making the model insensitive to perturbations can also suppress task-relevant signals needed to fit and generalize. Input-dependent routing---assigning different computation paths to different inputs---can relax this tension, but brings new fragility: distribution shifts may mislead routing decisions, and perturbations can destabilize routing, compounding downstream errors. We formalize these effects through two risk decompositions that separate (i) how well the available paths cover diverse test conditions from how accurately the router selects among them, and (ii) how sensitive each fixed path is from how much routing fluctuation amplifies that sensitivity. Guided by these analyses, we propose STEM-GNN: Stable TokEnized Mixture-of-Experts GNN, a pretrain-then-finetune framework that couples a mixture-of-experts encoder providing diverse computation paths to cover heterogeneous test conditions, a vector-quantized token interface that maps encoder outputs to a discrete codebook to absorb small representation drift induced by input perturbations and routing fluctuations before downstream layers, and a Lipschitz-regularized prediction head that bounds how much the output can amplify residual upstream variation. Across eight node, link, and graph benchmarks, STEM-GNN maintains strong clean performance; on representative node benchmarks, it improves the three-way balance under degree/homophily shifts and feature/edge perturbations. The code and data are available at https://github.com/GXG-CS/STEM-GNN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on33
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng et al.NeurIPS 2021 · 1,632 citations
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
Related papers
- Mitigating Dynamic Graph Distribution Shifts via Mixture of Variational ExpertsQianyu Song, Chao Li, Yeyu Yan, Hui Zhou et al.WWW 2026
- Discriminative Mixture-of-Experts on Graphs with Reliable Expert FusionHaoyue Deng, Menghui Wang, Yunlong Zhou, Jingyi Liu et al.ICML 2026
- Chasing All-Round Graph Representation Robustness: Model, Training, and OptimizationChunhui Zhang, Yijun Tian, Mingxuan Ju, Zheyuan Liu et al.ICLR 2023
- Graph Out-of-Distribution Generalization via Causal InterventionQitian Wu, Fan Nie, Chenxiao Yang, Tianyi Bao et al.WWW 2024 · 58 citations
- Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingHaotao Wang, Ziyu Jiang, Yuning You, Yan Han et al.NeurIPS 2023 · 104 citations
