Multi-Scale Distillation from Multiple Graph Neural Networks
Chunhai Zhang, Jie Liu, Kai Dang, Wenzheng Zhang
Abstract
Knowledge Distillation (KD), which is an effective model compression and acceleration technique, has been successfully applied to graph neural networks (GNNs) recently. Existing approaches utilize a single GNN model as the teacher to distill knowledge. However, we notice that GNN models with different number of layers demonstrate different classification abilities on nodes with different degrees. On the one hand, for nodes with high degrees, their local structures are dense and complex, hence more message passing is needed. Therefore, GNN models with more layers perform better. On the other hand, for nodes with low degrees, whose local structures are relatively sparse and simple, the repeated message passing can easily lead to over-smoothing. Thus, GNN models with less layers are more suitable. However, existing single-teacher GNN knowledge distillation approaches which are based on a single GNN model, are sub-optimal. To this end, we propose a novel approach to distill multi-scale knowledge, which learns from multiple GNN teacher models with different number of layers to capture the topological semantic at different scales. Instead of learning from the teacher models equally, the proposed method automatically assigns proper weights for each teacher model via an attention mechanism which enables the student to select teachers for different local structures. Extensive experiments are conducted to evaluate the proposed method on four public datasets. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods. Our code is publicly available at https://github.com/NKU-IIPLab/MSKD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e402c1e7-ad2b-491f-8f1b-75540af6109fCited by top-tier papers4
- Preference-driven Knowledge Distillation for Few-shot Node ClassificationXing Wei, Chunchun Chen, Rui Fan, Xiaofeng Cao et al.NeurIPS 2025 · 2 citations
- Geometry Awakening: Cross-Geometry Learning Exhibits Superiority over Individual StructuresYadong Sun, Xiaofeng Cao, Yu Wang, Wei Ye et al.NeurIPS 2024 · 2 citations
- Pre-Training Graph Contrastive Masked Autoencoders are Strong Distillers for EEGXinxu Wei, Kanhao Zhao, Yong Jiao, Hua Xie et al.ICML 2025
- Pareto-Based Heterogeneous Knowledge Distillation for MLPs on GraphsWenrui Zhao, Yijun Tian, Zhichao Xu, Yawei Wang et al.AAAI 2026
Builds on5
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang et al.AAAI 2021 · 368 citations
- Graph Few-Shot Learning via Knowledge TransferHuaxiu Yao, Chuxu Zhang, Ying Wei, Meng Jiang et al.AAAI 2020 · 193 citations
- Distilling Holistic Knowledge with Graph Neural NetworksSheng Zhou, Yucheng Wang, Defang Chen, Jiawei Chen et al.ICCV 2021 · 69 citations
- Amalgamating Knowledge From Heterogeneous Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.CVPR 2021
- Distilling Knowledge From Graph Convolutional NetworksYiding Yang, Jiayan Qiu, Mingli Song, Dacheng Tao et al.CVPR 2020
Related papers
- Boosting Graph Neural Networks via Adaptive Knowledge DistillationZhichun Guo, Chunhui Zhang, Yujie Fan, Yijun Tian et al.AAAI 2023 · 48 citations
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 28 citations
- Compressing Deep Graph Neural Networks via Adversarial Knowledge DistillationHuarui He, Jie Wang, Zhanqiu Zhang, Feng WuKDD 2022 · 44 citations
- MuGSI: Distilling GNNs with Multi-Granularity Structural Information for Graph ClassificationTianjun Yao, Jiaqi Sun, Defu Cao, Kun Zhang et al.WWW 2024 · 9 citations
- MARCH: Multi-Teacher Contrastive Hypergraph DistillationRongwei Xu, Zitai Qiu, Pengfei Ding, Jia Wu et al.WWW 2026
