Pareto-Based Heterogeneous Knowledge Distillation for MLPs on Graphs
Wenrui Zhao, Yijun Tian, Zhichao Xu, Yawei Wang, Chuxu Zhang
Abstract
Heterogeneous Graph Neural Networks (HGNNs) have demonstrated remarkable capabilities in capturing effective information in heterogeneous graphs, achieving outstanding performance in various learning tasks. However, their heavy dependency on neighbor information may result in high latency, which restricts their practicality in real-world applications. Recent studies have attempted to overcome such latency in Graph Neural Networks (GNNs) by distilling knowledge into student models that do not rely on graph structure. But these approaches primarily focus on replicating teachers' predictive outcomes while neglecting the structural knowledge they encoded. This limitation makes such approaches less effective when graphs become complex, particularly in heterogeneous graphs. Motivated by this challenge, we propose HGKD, a novel hierarchical knowledge distillation framework that transfers both structural knowledge and predictive outcomes from HGNN teachers to a multi-layer perceptron (MLP) student. Additionally, we provide two variants of HGKD that help the student learn from multiple teacher models via Pareto learning, and incorporate low-cost neighbor information. We evaluate HGKD and its variants on a range of heterogeneous graph datasets. The results demonstrate that the student model achieves performance comparable to, or even exceeding, that of HGNN teachers, despite not relying on graph structures during inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 56d1cb71-ee68-4e97-9a33-a41118f56acbBuilds on16
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- MAGNN: Metapath Aggregated Graph Neural Network for Heterogeneous Graph EmbeddingXinyu Fu, Jiani Zhang, Ziqiao Meng, Irwin KingWWW 2020 · 1,149 citations
- Graph-less Neural Networks: Teaching Old MLPs New Tricks Via DistillationShichang Zhang, Yozen Liu, Yizhou Sun, Neil ShahICLR 2022 · 234 citations
- Extract the Knowledge of Graph Neural Networks and Go Beyond it: An Effective Knowledge Distillation FrameworkCheng Yang, Jiawei Liu, Chuan ShiWWW 2021 · 153 citations
Related papers
- Boosting Graph Neural Networks via Adaptive Knowledge DistillationZhichun Guo, Chunhui Zhang, Yujie Fan, Yijun Tian et al.AAAI 2023 · 48 citations
- LightHGNN: Distilling Hypergraph Neural Networks into MLPs for 100x Faster InferenceYifan Feng, Yihe Luo, Shihui Ying, Yue GaoICLR 2024 · 8 citations
- Linkless Link Prediction via Relational DistillationZhichun Guo, William Shiao, Shichang Zhang, Yozen Liu et al.ICML 2023 · 60 citations
- Multi-Scale Distillation from Multiple Graph Neural NetworksChunhai Zhang, Jie Liu, Kai Dang, Wenzheng ZhangAAAI 2022 · 17 citations
- Quantifying the Knowledge in GNNs for Reliable Distillation into MLPsLirong Wu, Haitao Lin, Yufei Huang, Stan Z. LiICML 2023 · 48 citations
