Scaffolding a Student to Instill Knowledge
Anil Kag, Durmus Alp Emre Acar, Aditya Gangrade, Venkatesh Saligrama
Abstract
We propose a novel knowledge distillation (KD) method to selectively instill teacher knowledge into a student model motivated by situations where the student's capacity is significantly smaller than that of the teachers. In vanilla KD, the teacher primarily sets a predictive target for the student to follow, and we posit that this target is overly optimistic due to the student's lack of capacity. We develop a novel scaffolding scheme where the teacher, in addition to setting a predictive target, also scaffolds the student's prediction by censoring hard-to-learn examples. The student model utilizes the same information as the teacher's soft-max predictions as inputs, and in this sense, our proposal can be viewed as a natural variant of vanilla KD. We show on synthetic examples that censoring hard-examples leads to smoothening the student's loss landscape so that the student encounters fewer local minima. As a result, it has good generalization properties. Against vanilla KD, we achieve improved performance and are comparable to more intrusive techniques that leverage feature matching on benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang et al.AAAI 2021 · 368 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva et al.CVPR 2022 · 215 citations
Related papers
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 7 citations
- Hard Gate Knowledge Distillation - Leverage Calibration for Robust and Reliable Language ModelDongkyu Lee, Zhiliang Tian, Yingxiu Zhao, Ka Chun Cheung et al.EMNLP 2022
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
- Revisiting Knowledge Distillation: An Inheritance and Exploration FrameworkZhen Huang, Xu Shen, Jun Xing, Tongliang Liu et al.CVPR 2021
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok et al.EMNLP 2022 · 4 citations
