Learning an Inference-accelerated Network from a Pre-trained Model with Frequency-enhanced Feature Distillation
Xuesong Niu, Jili Gu, Guoxin Zhang, Pengfei Wan, Zhongyuan Wang
摘要
Convolution neural networks (CNNs) have achieved great success in various computer vision tasks, but they are still suffering from the heavy computation costs, which are mainly resulted from the substantial redundancy of the feature maps. In order to reduce these redundancy, we proposed a simple but effective frequency-enhanced feature distillation strategy to train an inference-accelerated network with a pre-trained model. Traditionally, one CNN can be regarded as a hierarchical structure, which can generate the low-level, middle-level and high-level feature maps from different convolution layers. In order to accelerate the inference time of CNNs, in this paper, we propose to resize the low-level and middle-level feature maps to smaller scales to reduce the spatial computation costs of CNNs. A frequency-enhanced feature distillation training strategy with a pre-trained model is then used to help the inference-accelerated network to maintain the core information after resizing the feature maps. To be specific, the original pre-trained network and the inference-accelerated network with resized feature maps are regarded as the teacher network and student network respectively. Considering that the low-frequency domain of the feature maps contribute the most parts to the final classification, we then transform the feature maps of different levels into a frequency-enhanced feature space, which highlights the low-frequency features for both the teacher and student networks. The frequency-enhanced features are used to transfer the knowledge from the teacher network to the student network. At the same time, knowledge for the final classification, i.e., the classification feature and predicted probabilities, are also used for distillation. Experiments on multiple databases based on various network structure types, e.g., ResNet, Res2Net, MobileNetV2, and ConvNeXt, have shown that with the proposed frequency-enhanced feature distillation training strategy, our method could get an inference-accelerated network with comparable performance and much less computation cost.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Ultrafast Video Attention Prediction with Coupled Knowledge DistillationKui Fu, Peipei Shi, Yafei Song, Shiming Ge 等AAAI 2020 · 被引用 11 次
- Efficient Crowd Counting via Structured Knowledge TransferLingbo Liu, Jiaqi Chen, Hefeng Wu, Tianshui Chen 等ACM MM 2020 · 被引用 72 次
- Pay Attention to Features, Transfer Learn Faster CNNsKafeng Wang, Xitong Gao, Yiren Zhao, Xingjian Li 等ICLR 2020 · 被引用 83 次
- Task Decoupled Knowledge Distillation For Lightweight Face DetectorsXiaoqing Liang, Xu Zhao, Chaoyang Zhao, Nanfei Jiang 等ACM MM 2020 · 被引用 5 次
