Feature-map-level Online Adversarial Knowledge Distillation
Inseop Chung, Seonguk Park, Jangho Kim, Nojun Kwak
Abstract
Feature maps contain rich information about image intensity and spatial correlation. However, previous online knowledge distillation methods only utilize the class probabilities. Thus in this paper, we propose an online knowledge distillation method that transfers not only the knowledge of the class probabilities but also that of the feature map using the adversarial training framework. We train multiple networks simultaneously by employing discriminators to distinguish the feature map distributions of different networks. Each network has its corresponding discriminator which discriminates the feature map from its own as fake while classifying that of the other network as real. By training a network to fool the corresponding discriminator, it can learn the other network's feature map distribution. We show that our method performs better than the conventional direct alignment method such as L 1 and is more suitable for online distillation. Also, we propose a novel cyclic learning scheme for training more than two networks together. We have applied our method to various network architectures on the classification task and discovered a significant improvement of performance especially in the case of training a pair of a small network and a large one.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian SongNeurIPS 2022 · 88 citations
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial EstimationAlexandre Ramé, Matthieu CordICLR 2021 · 60 citations
- Compressing Deep Graph Neural Networks via Adversarial Knowledge DistillationHuarui He, Jie Wang, Zhanqiu Zhang, Feng WuKDD 2022 · 44 citations
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 19 citations
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen et al.AAAI 2023 · 16 citations
Related papers
- Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge DistillationMengxue Kang, Jinpeng Zhang, Jinming Zhang, Xiashuang Wang et al.ICCV 2023 · 23 citations
- Towards Learning Spatially Discriminative Feature RepresentationsChaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang et al.ICCV 2021 · 23 citations
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu et al.CVPR 2020
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng et al.ICCV 2023 · 16 citations
- Discriminator-Cooperated Feature Map Distillation for GAN CompressionTie Hu, Mingbao Lin, Lizhou You, Fei Chao et al.CVPR 2023
