Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
Chengming Hu, Haolun Wu, Xuan Li, Chen Ma, Xi Chen, Boyu Wang, Jun Yan, Xue Liu
Abstract
Knowledge distillation aims to train a compact student network using soft supervision from a larger teacher network and hard supervision from ground truths. However, determining an optimal knowledge fusion ratio that balances these supervisory signals remains challenging. Prior methods generally resort to a constant or heuristic-based fusion ratio, which often falls short of a proper balance. In this study, we introduce a novel adaptive method for learning a sample-wise knowledge fusion ratio, exploiting both the correctness of teacher and student, as well as how well the student mimics the teacher on each sample. Our method naturally leads to the intra-sample trilateral geometric relations among the student prediction (S), teacher prediction (T ), and ground truth (G). To counterbalance the impact of outliers, we further extend to the inter-sample relations, incorporating the teacher's global average prediction ( T ) for samples within the same class. A simple neural network then learns the implicit mapping from the intra-and inter-sample relations to an adaptive, sample-wise knowledge fusion ratio in a bilevel-optimization manner. Our approach provides a simple, practical, and adaptable solution for knowledge distillation that can be employed across various architectures and model sizes. Extensive experiments demonstrate consistent improvements over other loss re-weighting methods on image classification, attack detection, and click-through rate prediction. * Equal contribution with random order.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2e3ad3a-52ae-4bfd-97d9-37ad4b950786Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou et al.ICLR 2021 · 209 citations
- Out-of-Distribution Detection via Conditional Kernel Independence ModelYu Wang, Jingjing Zou, Jingyang Lin, Qing Ling et al.NeurIPS 2022 · 9 citations
Related papers
- Uncertainty-Aware Knowledge Distillation for Multimodal Large Language ModelsJingchen Sun, Shaobo Han, Deep Patel, Wataru Kohno et al.CVPR 2026 · 2 citations
- Rethinking the Dark Knowledge and Kullback-Leibler Divergence Loss in Knowledge Distillation Under Capacity MismatchingYingchao Wang, Wenqi Niu, Xingshan Yao, Li You et al.AAAI 2026
- EA-KD: Entropy-Based Adaptive Knowledge DistillationChi-Ping Su, Ching-Hsun Tseng, Bin Pu, Lei Zhao et al.ICCV 2025 · 5 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 4 citations
