Distilling Holistic Knowledge with Graph Neural Networks
Sheng Zhou, Yucheng Wang, Defang Chen, Jiawei Chen, Xin Wang, Can Wang, Jiajun Bu
Abstract
Knowledge Distillation (KD) aims at transferring knowledge from a larger well-optimized teacher network to a smaller learnable student network. Existing KD methods have mainly considered two types of knowledge, namely the individual knowledge and the relational knowledge. However, these two types of knowledge are usually modeled independently while the inherent correlations between them are largely ignored. It is critical for sufficient student network learning to integrate both individual knowledge and relational knowledge while reserving their inherent correlation. In this paper, we propose to distill the novel holistic knowledge based on an attributed graph constructed among instances. The holistic knowledge is represented as a unified graph-based embedding by aggregating individual knowledge from relational neighborhood samples with graph neural networks, the student network is learned by distilling the holistic knowledge in a contrastive manner. Extensive experiments and ablation studies are conducted on benchmark datasets, the results demonstrate the effectiveness of the proposed method. The code has been published in https://github.com/wyc-ruiker/HKD
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- Collaborative Knowledge Distillation for Heterogeneous Information Network EmbeddingCan Wang, Sheng Zhou, Kang Yu, Defang Chen et al.WWW 2022 · 46 citations
- Patch-Wise Graph Contrastive Learning for Image TranslationChanyong Jung, Gihyun Kwon, Jong Chul YeAAAI 2024 · 22 citations
- Partition Speeds Up Learning Implicit Neural Representations Based on Exponential-Increase HypothesisKe Liu, Feng Liu, Haishuai Wang, Ning Ma et al.ICCV 2023 · 19 citations
- Bending Graphs: Hierarchical Shape Matching using Gated Optimal TransportMahdi Saleh, Shun-Cheng Wu, Luca Cosmo, Nassir Navab et al.CVPR 2022 · 17 citations
Builds on6
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang et al.AAAI 2021 · 368 citations
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
Related papers
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu et al.CVPR 2021
- Do Topological Characteristics Help in Knowledge Distillation?Jungeun Kim, Junwon You, Dongjin Lee, Ha Young Kim et al.ICML 2024 · 11 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng et al.ICCV 2023 · 16 citations
- VRM: Knowledge Distillation via Virtual Relation MatchingWeijia Zhang, Fei Xie, Tom Weidong Cai, Chao MaICCV 2025 · 6 citations
