Distilling Cross-Task Knowledge via Relationship Matching
Han-Jia Ye, Su Lu, De-Chuan Zhan
Abstract
The discriminative knowledge from a high-capacity deep neural network (a.k.a. the "teacher") could be distilled to facilitate the learning efficacy of a shallow counterpart (a.k.a. the "student"). This paper deals with a general scenario reusing the knowledge from a cross-task teachertwo models are targeting non-overlapping label spaces. We emphasize that the comparison ability between instances acts as an essential factor threading knowledge across domains, and propose the RElationship FacIlitated Local cLassifiEr Distillation (REFILLED) approach, which decomposes the knowledge distillation flow into branches for embedding and the top-layer classifier. In particular, different from reconciling the instance-label confidence between models, REFILLED requires the teacher to reweight the hard triplets push forwarded by the student so that the similarity comparison levels between instances are matched. A local embedding-induced classifier from the teacher further supervises the student's classification confidence. RE-FILLED demonstrates its effectiveness when reusing crosstask models, and also achieves state-of-the-art performance on the standard knowledge distillation benchmarks. The code of the paper can be accessed at https://github . com/njulus/ReFilled.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bfe4ab4-fd63-45bf-aa4c-62c6424a1673Cited by top-tier papers7
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li et al.NeurIPS 2022 · 46 citations
- Tailoring Embedding Function to Heterogeneous Few-Shot Tasks by Global and Local Feature AdaptorsSu Lu, Han-Jia Ye, De-Chuan ZhanAAAI 2021 · 29 citations
- Task Cooperation for Semi-Supervised Few-Shot LearningHan-Jia Ye, Xin-Chun Li, De-Chuan ZhanAAAI 2021 · 20 citations
- MLink: Linking Black-Box Models for Collaborative Multi-Model InferenceMu Yuan, Lan Zhang, Xiang-Yang LiAAAI 2022 · 10 citations
- Cross-domain Knowledge Distillation for Retrieval-based Question Answering SystemsCen Chen, Chengyu Wang, Minghui Qiu, Dehong Gao et al.WWW 2021 · 10 citations
Builds on4
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- UM-Adapt: Unsupervised Multi-Task Adaptation Using Adversarial Cross-Task DistillationJogendra Nath Kundu, Nishank Lakkakula, Venkatesh Babu RadhakrishnanICCV 2019 · 62 citations
Related papers
- Distilling Knowledge via Knowledge ReviewPengguang Chen, Shu Liu, Hengshuang Zhao, Jiaya JiaCVPR 2021
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 10 citations
- Knowledge Refinery: Learning from Decoupled LabelQianggang Ding, Sifan Wu, Tao Dai, Hao Sun et al.AAAI 2021 · 15 citations
- Exploring the Knowledge Transferred by Response-Based Teacher-Student DistillationLiangchen Song, Xuan Gong, Helong Zhou, Jiajie Chen et al.ACM MM 2023 · 14 citations
- Joint Pre-training and Local Re-training: Transferable Representation Learning on Multi-source Knowledge GraphsZequn Sun, Jiacheng Huang, Jinghao Lin, Xiaozhou Xu et al.KDD 2023 · 5 citations
