VRM: Knowledge Distillation via Virtual Relation Matching
Weijia Zhang, Fei Xie, Tom Weidong Cai, Chao Ma
Abstract
Knowledge distillation (KD) aims to transfer the knowledge of a more capable yet cumbersome teacher model to a lightweight student model. In recent years, relation-based KD methods have fallen behind, as their instance-matching counterparts dominate in performance. In this paper, we revive relational KD by identifying and tackling several key issues in relation-based methods, including their susceptibility to overfitting and spurious responses. Specifically, we transfer novelly constructed affinity graphs that compactly encapsulate a wealth of beneficial inter-sample, inter-class, and inter-view correlations by exploiting virtual views and relations as a new kind of knowledge. As a result, the student has access to richer guidance signals and stronger regularisation throughout the distillation process. To further mitigate the adverse impact of spurious responses, we prune the affinity graphs by dynamically detaching redundant and unreliable edges. Extensive experiments on CIFAR-100, ImageNet, and MS-COCO datasets demonstrate the superior performance of the proposed virtual relation matching (VRM) method, where it consistently sets new state-of-theart records over a range of models, architectures, tasks, and set-ups. For instance, VRM for the first time hits 74.0 % accuracy for ResNet50 → MobileNetV2 distillation on ImageNet, and improves DeiT-T by 14.44 % on CIFAR-100 with a ResNet56 teacher. The code and models are released at https://github.com/VISION-SJTU/VRM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9314dec-c8e7-4edb-aa31-325458cfe816Cited by top-tier papers1
Ask how each one uses itBuilds on47
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- General Instance Distillation for Object DetectionXing Dai, Zeren Jiang, Zhao Wu, Yiping Bao et al.CVPR 2021
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
- Exploring Inter-Channel Correlation for Diversity-preserved Knowledge DistillationLi Liu, Qingle Huang, Sihao Lin, Hongwei Xie et al.ICCV 2021 · 132 citations
- Inter-Region Affinity Distillation for Road Marking SegmentationYuenan Hou, Zheng Ma, Chunxiao Liu, Tak-Wai Hui et al.CVPR 2020
