Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model
Liangchen Song, Jialian Wu, Ming Yang, Qian Zhang, Yuan Li, Junsong Yuan
摘要
When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf models is not trivial and generally requires considerable try-and-error and parameter tuning. In this paper, we denote a well-trained model as a teacher network and a model for the new task as a student network. We aim to ease the efforts of transferring knowledge from the teacher to the student network, robust to the gaps between their network architectures, domain data, and task definitions. Specifically, we propose a hybrid forward scheme in training the teacher-student models, alternately updating layer weights of the student model. The key merit of our hybrid forward scheme is on the dynamical balance between the knowledge transfer loss and task specific loss in training. We demonstrate the effectiveness of our method on a variety of tasks, e.g., model compression, segmentation, and detection, under a variety of knowledge transfer settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Ensemble Attention Distillation for Privacy-Preserving Federated LearningXuan Gong, Abhishek Sharma, Srikrishna Karanam, Ziyan Wu 等ICCV 2021 · 被引用 148 次
- Handling Difficult Labels for Multi-label Image Classification via Uncertainty DistillationLiangchen Song, Jialian Wu, Ming Yang, Qian Zhang 等ACM MM 2021 · 被引用 15 次
它引用的顶会 Paper10
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei 等EMNLP 2020 · 被引用 168 次
相关 Paper
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 被引用 10 次
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 被引用 103 次
- Hierarchical Knowledge Squeezed Adversarial Network CompressionPeng Li, Chang Shu, Yuan Xie, Yan Qu 等AAAI 2020 · 被引用 6 次
- Distilling Knowledge via Knowledge ReviewPengguang Chen, Shu Liu, Hengshuang Zhao, Jiaya JiaCVPR 2021
- Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge AmalgamationChengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song 等ICCV 2019 · 被引用 63 次
