Diverse and Sparse Mixture-of-Experts for Causal Subgraph-Based Out-of-Distribution Graph Learning
Jerry Sun, Mohamed Abubakr Hassan, Yaoyu Zhang, Wanying Zhang, Chi-Guhn Lee
Abstract
Current state-of-the-art methods for out-of-distribution (OOD) generalization lack the ability to effectively address datasets with heterogeneous causal subgraphs at the instance level. Existing approaches that attempt to handle such heterogeneity either rely on data augmentation, which risks altering label semantics, or impose causal assumptions whose validity in real-world datasets is uncertain. We introduce DiSCO, a novel Mixture-of-Experts (MoE) framework for Diversity- and Sparsity-driven Causal OOD graph learning, designed to model heterogeneous causal subgraphs without relying on restrictive assumptions. Our key idea is to address instance-level heterogeneity by enforcing semantic diversity among experts, each generating a distinct causal subgraph, while a learned gate assigns sparse weights that adaptively focus on the most relevant experts for each input. Our theoretical analysis shows that these two properties jointly reduce OOD error. In practice, our experts are scalable and do not require environment labels. Empirically, our framework achieves strong performance on the GOOD benchmark across both synthetic and real-world structural shifts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd417b9f-587e-4fa6-848f-b0c5db439befBuilds on15
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Interpretable and Generalizable Graph Learning via Stochastic Attention MechanismSiqi Miao, Mia Liu, Pan LiICML 2022 · 288 citations
- Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsYongqiang Chen, Yonggang Zhang, Yatao Bian, Han Yang et al.NeurIPS 2022 · 246 citations
Related papers
- Invariant Learning on Heterogeneous Graphs via Subgraph Environment InferenceYanghui Fu, Yunfei Wang, Hao Zou, Yue He et al.WWW 2026
- Joint Learning of Label and Environment Causal Independence for Graph Out-of-Distribution GeneralizationShurui Gui, Meng Liu, Xiner Li, Youzhi Luo et al.NeurIPS 2023 · 54 citations
- Improving Out-of-Distribution Generalization in Graphs via Hierarchical Semantic EnvironmentsYinhua Piao, Sangseon Lee, Yijingxiu Lu, Sun KimCVPR 2024 · 6 citations
- Quantifying Distributional Invariance in Causal Subgraph for IRM-Free Graph GeneralizationYang Qiu, Yixiong Zou, Jun Wang, Wei Liu et al.NeurIPS 2025 · 3 citations
- Improving Generalization of Dynamic Graph Learning via Environment PromptKuo Yang, Zhengyang Zhou, Qihe Huang, Limin Li et al.NeurIPS 2024 · 14 citations
