Generalizing GNNs with Tokenized Mixture of Experts
Xiaoguang Guo, Zehong Wang, Jiazheng Li, Shawn Spitzel, Qi Yang, Kaize Ding, Jundong Li, Chuxu Zhang
摘要
Deployed graph neural networks (GNNs) operate as frozen snapshots, yet must simultaneously fit clean data, generalize under distribution shifts, and remain stable against input perturbations---three goals that are difficult to satisfy at once with a single fixed model. We first show theoretically that a single fixed inference rule can create a stability--generalization tradeoff: making the model insensitive to perturbations can also suppress task-relevant signals needed to fit and generalize. Input-dependent routing---assigning different computation paths to different inputs---can relax this tension, but brings new fragility: distribution shifts may mislead routing decisions, and perturbations can destabilize routing, compounding downstream errors. We formalize these effects through two risk decompositions that separate (i) how well the available paths cover diverse test conditions from how accurately the router selects among them, and (ii) how sensitive each fixed path is from how much routing fluctuation amplifies that sensitivity. Guided by these analyses, we propose STEM-GNN: Stable TokEnized Mixture-of-Experts GNN, a pretrain-then-finetune framework that couples a mixture-of-experts encoder providing diverse computation paths to cover heterogeneous test conditions, a vector-quantized token interface that maps encoder outputs to a discrete codebook to absorb small representation drift induced by input perturbations and routing fluctuations before downstream layers, and a Lipschitz-regularized prediction head that bounds how much the output can amplify residual upstream variation. Across eight node, link, and graph benchmarks, STEM-GNN maintains strong clean performance; on representative node benchmarks, it improves the three-way balance under degree/homophily shifts and feature/edge perturbations. The code and data are available at https://github.com/GXG-CS/STEM-GNN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 被引用 1,599 次
相关 Paper
- Mitigating Dynamic Graph Distribution Shifts via Mixture of Variational ExpertsQianyu Song, Chao Li, Yeyu Yan, Hui Zhou 等WWW 2026
- Discriminative Mixture-of-Experts on Graphs with Reliable Expert FusionHaoyue Deng, Menghui Wang, Yunlong Zhou, Jingyi Liu 等ICML 2026
- Chasing All-Round Graph Representation Robustness: Model, Training, and OptimizationChunhui Zhang, Yijun Tian, Mingxuan Ju, Zheyuan Liu 等ICLR 2023
- Graph Out-of-Distribution Generalization via Causal InterventionQitian Wu, Fan Nie, Chenxiao Yang, Tianyi Bao 等WWW 2024 · 被引用 58 次
- Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingHaotao Wang, Ziyu Jiang, Yuning You, Yan Han 等NeurIPS 2023 · 被引用 104 次
