Feature-map-level Online Adversarial Knowledge Distillation
Inseop Chung, Seonguk Park, Jangho Kim, Nojun Kwak
摘要
Feature maps contain rich information about image intensity and spatial correlation. However, previous online knowledge distillation methods only utilize the class probabilities. Thus in this paper, we propose an online knowledge distillation method that transfers not only the knowledge of the class probabilities but also that of the feature map using the adversarial training framework. We train multiple networks simultaneously by employing discriminators to distinguish the feature map distributions of different networks. Each network has its corresponding discriminator which discriminates the feature map from its own as fake while classifying that of the other network as real. By training a network to fool the corresponding discriminator, it can learn the other network's feature map distribution. We show that our method performs better than the conventional direct alignment method such as L 1 and is more suitable for online distillation. Also, we propose a novel cyclic learning scheme for training more than two networks together. We have applied our method to various network architectures on the classification task and discovered a significant improvement of performance especially in the case of training a pair of a small network and a large one.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Bi-directional Weakly Supervised Knowledge Distillation for Whole Slide Image ClassificationLinhao Qu, Xiaoyuan Luo, Manning Wang, Zhijian SongNeurIPS 2022 · 被引用 88 次
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial EstimationAlexandre Ramé, Matthieu CordICLR 2021 · 被引用 60 次
- Compressing Deep Graph Neural Networks via Adversarial Knowledge DistillationHuarui He, Jie Wang, Zhanqiu Zhang, Feng WuKDD 2022 · 被引用 44 次
- NORM: Knowledge Distillation via N-to-One Representation MatchingXiaolong Liu, Lujun Li, Chao Li, Anbang YaoICLR 2023 · 被引用 19 次
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen 等AAAI 2023 · 被引用 16 次
相关 Paper
- Alleviating Catastrophic Forgetting of Incremental Object Detection via Within-Class and Between-Class Knowledge DistillationMengxue Kang, Jinpeng Zhang, Jinming Zhang, Xiashuang Wang 等ICCV 2023 · 被引用 23 次
- Towards Learning Spatially Discriminative Feature RepresentationsChaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang 等ICCV 2021 · 被引用 23 次
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu 等CVPR 2020
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng 等ICCV 2023 · 被引用 16 次
- Discriminator-Cooperated Feature Map Distillation for GAN CompressionTie Hu, Mingbao Lin, Lizhou You, Fei Chao 等CVPR 2023
