Customizing Student Networks From Heterogeneous Teachers via Adaptive Knowledge Amalgamation
Chengchao Shen, Mengqi Xue, Xinchao Wang, Jie Song, Li Sun, Mingli Song
Abstract
A massive number of well-trained deep networks have been released by developers online. These networks may focus on different tasks and in many cases are optimized for different datasets. In this paper, we study how to exploit such heterogeneous pre-trained networks, known as teachers, so as to train a customized student network that tackles a set of selective tasks defined by the user. We assume no human annotations are available, and each teacher may be either single- or multi-task. To this end, we introduce a dual-step strategy that first extracts the task-specific knowledge from the heterogeneous teachers sharing the same sub-task, and then amalgamates the extracted knowledge to build the student network. To facilitate the training, we employ a selective learning scheme where, for each unlabelled sample, the student learns adaptively from only the teacher with the least prediction ambiguity. We evaluate the proposed approach on several datasets and the experimental results demonstrate that the student, learned by such adaptive knowledge amalgamation, achieves performances even better than those of the teachers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68b59d03-08ff-4e35-a0c9-4ee4f17c55ecCited by top-tier papers24
- One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge DistillationZhiwei Hao, Jianyuan Guo, Kai Han, Yehui Tang et al.NeurIPS 2023 · 205 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song et al.AAAI 2021 · 55 citations
- Mosaicking to Distill: Knowledge Distillation from Out-of-Domain DataGongfan Fang, Yifan Bao, Jie Song, Xinchao Wang et al.NeurIPS 2021 · 53 citations
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.ICCV 2021 · 46 citations
Builds on1
Related papers
- Data-Free Knowledge Amalgamation via Group-Stack Dual-GANJingwen Ye, Yixin Ji, Xinchao Wang, Xin Gao et al.CVPR 2020
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Amalgamating Knowledge From Heterogeneous Graph Neural NetworksYongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song et al.CVPR 2021
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim et al.NeurIPS 2021 · 134 citations
- Knowledge Amalgamation for Multi-Label Classification via Label Dependency TransferJidapa Thadajarassiri, Thomas Hartvigsen, Walter Gerych, Xiangnan Kong et al.AAAI 2023 · 9 citations
