Discriminative Mixture-of-Experts on Graphs with Reliable Expert Fusion
Haoyue Deng, Menghui Wang, Yunlong Zhou, Jingyi Liu, Ran Zhang, Chunming Hu, Xiao Wang
Abstract
Graph Mixture-of-Experts (Graph-MoE) offers a way to scale GNNs via adaptive capacity allocation, with the goal of allowing different experts to capture diverse graph patterns. Its effectiveness heavily depends on the coordination between routing decisions and expert specialization. However, through extensive empirical study, we identify two critical phenomena. First, discrimination loss occurs on both the expert and routing sides, where GNN experts become highly homogenized and the router collapses to a small subset of experts, failing to reflect diverse graph semantics. Second, routing uncertainty is prevalent, as existing routers produce uncertain expert assignments for most nodes, and such uncertainty exhibits a strong negative correlation with model performance. To address these issues, we propose C 2 GMoE, a novel Graph-MoE framework featuring Contrastive routing and Confidenceaware fusion. We introduce a group-wise contrastive routing strategy that provides explicit guidance for routing optimization by aligning node-level routing decisions with semantic clusters while satisfying load-balancing constraints. Moreover, through a theoretical analysis of generalization error, we develop a confidence-aware fusion mechanism that adaptively reweights expert predictions according to their confidence. Extensive experiments across multiple benchmarks demonstrate the effectiveness of our proposed C 2 GMoE. Code is publicly available at https: //github.com/Gatlin1111/C2GMoE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Graph Neural Networks Exponentially Lose Expressive Power for Node ClassificationKenta Oono, Taiji SuzukiICLR 2020 · 864 citations
- Beyond Low-frequency Information in Graph Convolutional NetworksDeyu Bo, Xiao Wang, Chuan Shi, Huawei ShenAAAI 2021 · 773 citations
Related papers
- One For All: Achieving Adaptive Graph Neural Networks via Mixture of Message PassingZhaojun Luo, Jintang Li, Yuchang Zhu, Yun Fu et al.KDD 2026
- Where Graph Meets Heterogeneity: Multi-View Collaborative Graph ExpertsZhihao Wu, Jinyu Cai, Yunhe Zhang, Jielong Lu et al.NeurIPS 2025 · 6 citations
- Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingHaotao Wang, Ziyu Jiang, Yuning You, Yan Han et al.NeurIPS 2023 · 104 citations
- C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts ModelsLin Li, Yan Wang, Zhuopeng WangAAAI 2026 · 1 citation
- Mixture of Weak and Strong Experts on GraphsHanqing Zeng, Hanjia Lyu, Diyi Hu, Yinglong Xia et al.ICLR 2024 · 11 citations
