Distilling Cross-Task Knowledge via Relationship Matching
Han-Jia Ye, Su Lu, De-Chuan Zhan
摘要
The discriminative knowledge from a high-capacity deep neural network (a.k.a. the "teacher") could be distilled to facilitate the learning efficacy of a shallow counterpart (a.k.a. the "student"). This paper deals with a general scenario reusing the knowledge from a cross-task teachertwo models are targeting non-overlapping label spaces. We emphasize that the comparison ability between instances acts as an essential factor threading knowledge across domains, and propose the RElationship FacIlitated Local cLassifiEr Distillation (REFILLED) approach, which decomposes the knowledge distillation flow into branches for embedding and the top-layer classifier. In particular, different from reconciling the instance-label confidence between models, REFILLED requires the teacher to reweight the hard triplets push forwarded by the student so that the similarity comparison levels between instances are matched. A local embedding-induced classifier from the teacher further supervises the student's classification confidence. RE-FILLED demonstrates its effectiveness when reusing crosstask models, and also achieves state-of-the-art performance on the standard knowledge distillation benchmarks. The code of the paper can be accessed at https://github . com/njulus/ReFilled.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
- Tailoring Embedding Function to Heterogeneous Few-Shot Tasks by Global and Local Feature AdaptorsSu Lu, Han-Jia Ye, De-Chuan ZhanAAAI 2021 · 被引用 29 次
- Task Cooperation for Semi-Supervised Few-Shot LearningHan-Jia Ye, Xin-Chun Li, De-Chuan ZhanAAAI 2021 · 被引用 20 次
- MLink: Linking Black-Box Models for Collaborative Multi-Model InferenceMu Yuan, Lan Zhang, Xiang-Yang LiAAAI 2022 · 被引用 10 次
- Cross-domain Knowledge Distillation for Retrieval-based Question Answering SystemsCen Chen, Chengyu Wang, Minghui Qiu, Dehong Gao 等WWW 2021 · 被引用 10 次
它引用的顶会 Paper4
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task DistillationJogendra Nath Kundu, Nishank Lakkakula, Venkatesh Babu RadhakrishnanICCV 2019 · 被引用 62 次
相关 Paper
- Distilling Knowledge via Knowledge ReviewPengguang Chen, Shu Liu, Hengshuang Zhao, Jiaya JiaCVPR 2021
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 被引用 10 次
- Knowledge Refinery: Learning from Decoupled LabelQianggang Ding, Sifan Wu, Tao Dai, Hao Sun 等AAAI 2021 · 被引用 15 次
- Exploring the Knowledge Transferred by Response-Based Teacher-Student DistillationLiangchen Song, Xuan Gong, Helong Zhou, Jiajie Chen 等ACM MM 2023 · 被引用 14 次
- Joint Pre-training and Local Re-training: Transferable Representation Learning on Multi-source Knowledge GraphsZequn Sun, Jiacheng Huang, Jinghao Lin, Xiaozhou Xu 等KDD 2023 · 被引用 5 次
