Narrow the Input Mismatch in Deep Graph Neural Network Distillation
Qiqi Zhou, Yanyan Shen, Lei Chen
Abstract
Graph neural networks (GNNs) have been widely studied for modeling graph-structured data. Thanks to the over-parameterization and large receptive field of deep GNNs, "deep" is a promising direction to develop GNNs further and has shown some superior performances. However, the over-stacked structures of deep architectures incur high inference cost in deployment. To compress deep GNNs, we can use knowledge distillation (KD) to make shallow student GNNs mimic teacher GNNs. Existing KD methods in graph domain focus on constructing diverse supervision on embedding or prediction produced by student GNNs, but overlook the gap of the receptive field (i.e., input information) between student and teacher, which brings difficulties to KD. We call this gap "input mismatch". To alleviate this problem, we propose a lightweight stochastic extended module to provide an estimation for missing input information for student GNNs. The estimator models the distribution of missing information. Specifically, we model the missing information as an independent distribution from graph level and a conditional distribution from node level (given the condition of observable input). These two estimates are optimized using a Bayesian methodology and combined into a balanced estimate as additional input to student GNNs. To the best of our knowledge, we are the first to address the "input mismatch" problem in deep GNNs distillation. Experiments on extensive benchmarks demonstrate that our method outperforms existing KD methods for GNNs in distillation performance, which confirms that the estimations are reasonable and effective.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f7375a68-7c86-499d-beec-aa2b6dae9820Cited by top-tier papers3
- Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference ServingShihong Gao, Xin Zhang, Yanyan Shen, Lei ChenSIGMOD 2025 · 7 citations
- Faster Convergence in Mini-batch Graph Neural Networks Training with Pseudo Full Neighborhood CompensationQiqi Zhou, Yanyan Shen, Lei ChenVLDB 2025 · 2 citations
- Efficient GNN Training on Giant Graphs with Collective Batching and SchedulingXin Zhang, Yanyan Shen, Yingxia Shao, Haoyang Li et al.VLDB 2026
Related papers
- Compressing Deep Graph Neural Networks via Adversarial Knowledge DistillationHuarui He, Jie Wang, Zhanqiu Zhang, Feng WuKDD 2022 · 44 citations
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 28 citations
- Boosting Graph Neural Networks via Adaptive Knowledge DistillationZhichun Guo, Chunhui Zhang, Yujie Fan, Yijun Tian et al.AAAI 2023 · 48 citations
- Multi-Scale Distillation from Multiple Graph Neural NetworksChunhai Zhang, Jie Liu, Kai Dang, Wenzheng ZhangAAAI 2022 · 17 citations
- T2-GNN: Graph Neural Networks for Graphs with Incomplete Features and Structure via Teacher-Student DistillationCuiying Huo, Di Jin, Yawen Li, Dongxiao He et al.AAAI 2023 · 74 citations
