Knowledge Distillation via Route Constrained Optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, Xiaolin Hu
摘要
Distillation-based learning boost the performance of the miniaturized neural network based on the hypothesis that the representation of a teacher model can be used as structured and relatively weak supervision, and thus would be easily learned by a miniaturized model. However, we find that the representation of a converged heavy model is still a strong constraint for training a small student model, which leads to a high lower bound of congruence loss. In this work, inspired by [1] we consider the knowledge distillation from the perspective of curriculum learning by routing. Instead of supervising the student model with a converged teacher model, we supervised it with some anchor points selected from the route in parameter space that the teacher model passed by, as we called route constrained optimization (RCO). We experimentally demonstrate this simple operation greatly reduces the lower bound of congruence loss for knowledge distillation, hint and mimicking learning. On close-set classification tasks like CIFAR [16] and ImageNet [3], RCO improves knowledge distillation by 2.14% and 1.5% respectively. For the sake of evaluating the generalization, we also test RCO on the open-set face recognition task MegaFace. RCO achieves 84.3% accuracy on 1 vs. 1 million task with only 0.8 M parameters, which push the SOTA by a large margin. * Equal contribution. research focus in recent years. Many methods were proposed to tackle this problem, such as model pruning [7, 17] , quantization [14, 29] and knowledge transfer [11, 25] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao 等AAAI 2023 · 被引用 277 次
- Densely Guided Knowledge Distillation using Multiple Teacher AssistantsWonchul Son, Jaemin Na, Junyong Choi, Wonjun HwangICCV 2021 · 被引用 158 次
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
- BERT Learns to Teach: Knowledge Distillation with Meta LearningWangchunshu Zhou, Canwen Xu, Julian J. McAuleyACL 2022 · 被引用 114 次
- Towards Efficient 3D Object Detection with Knowledge DistillationJihan Yang, Shaoshuai Shi, Runyu Ding, Zhe Wang 等NeurIPS 2022 · 被引用 76 次
相关 Paper
- DOT: A Distillation-Oriented TrainerBorui Zhao, Quan Cui, Renjie Song, Jiajun LiangICCV 2023 · 被引用 15 次
- Hierarchical Knowledge Squeezed Adversarial Network CompressionPeng Li, Chang Shu, Yuan Xie, Yan Qu 等AAAI 2020 · 被引用 6 次
- Evaluation-oriented Knowledge Distillation for Deep Face RecognitionYuge Huang, Jiaxiang Wu, Xingkun Xu, Shouhong DingCVPR 2022 · 被引用 35 次
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 被引用 48 次
- Unsupervised Representation Transfer for Small Networks: I Believe I Can Distill On-the-FlyHee Min Choi, Hyoa Kang, Dokwan OhNeurIPS 2021 · 被引用 13 次
