Demystifying GNN-to-MLP Knowledge Transfer: Theoretical Grounding and Dual-Stream Distillation Method
Zhiyuan Yu, Mingkai Lin, Wenzhong Li, Zhangyue Yin, Shijian Xiao, Sanglu Lu
Abstract
Graph Neural Networks (GNNs) have shown remarkable effectiveness across various applications, but their computational complexity poses significant scalability challenges. To this end, GNN-to-MLP Knowledge Distillation (KD) methods transfer relational inductive biases from GNNs to MLPs, equipping MLPs with graph-aware capabilities that rival or even surpass those of their teacher GNNs. However, a theoretical foundation for understanding GNN-to-MLP KD is still missing. In this paper, we provide a theoretical analysis of how knowledge distillation unlocks the potential of MLPs for graph tasks from the perspective of training dynamics. We demonstrate that label alignment in KD fundamentally reshapes the Neural Tangent Kernel (NTK) matrix of student MLPs, enabling them to learn the teacher model’s implicit graph bias. We further investigate finer-grained distillation paradigms and reveal that conventional layer-wise output alignment fails to effectively align the deep-layer graph propagation outcomes. To address this, we propose Dual-Stream Aligned MLP (DA-MLP), which incorporates complementary graph filters in a dual-stream architecture. This approach simultaneously enhances feature space dimensionality for improved representation alignment and preserves graph signals across different frequency bands. Comprehensive experiments on seven benchmark datasets validate that DA-MLP can be seamlessly integrated into existing knowledge distillation frameworks for performance enhancements in both transductive and inductive settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9c9f84b-7f83-4b85-80ea-1c1977bc65adBuilds on14
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- Graph Information BottleneckTailin Wu, Hongyu Ren, Pan Li, Jure LeskovecNeurIPS 2020 · 366 citations
- Revisiting Heterophily For Graph Neural NetworksSitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu et al.NeurIPS 2022 · 351 citations
- Graph-less Neural Networks: Teaching Old MLPs New Tricks Via DistillationShichang Zhang, Yozen Liu, Yizhou Sun, Neil ShahICLR 2022 · 234 citations
Related papers
- Extracting Low-/High- Frequency Knowledge from Graph Neural Networks and Injecting It into MLPs: An Effective GNN-to-MLP Distillation FrameworkLirong Wu, Haitao Lin, Yufei Huang, Tianyu Fan et al.AAAI 2023 · 52 citations
- TINED: GNNs-to-MLPs by Teacher Injection and Dirichlet Energy DistillationZiang Zhou, Zhihao Ding, Jieming Shi, Qing Li et al.ICML 2025
- AdaGMLP: AdaBoosting GNN-to-MLP Knowledge DistillationWeigang Lu, Ziyu Guan, Wei Zhao, Yaming YangKDD 2024 · 10 citations
- Leap of FAITH from GNN-to-MLP: Fairness Aware Inference via DisTillation of GrapH KnowledgeVipul Kumar Singh, Jyotismita Barman, Sandeep Kumar, Tapan K. Gandhi et al.AAAI 2026
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPsLing Yang, Ye Tian, Minkai Xu, Zhongyi Liu et al.ICLR 2024 · 48 citations
