Comprehensive Knowledge Distillation with Causal Intervention
Xiang Deng, Zhongfei Zhang
Abstract
Knowledge distillation (KD) addresses model compression by distilling knowledge from a large model (teacher) to a smaller one (student). The existing distillation approaches mainly focus on using different criteria to align the sample representations learned by the student and the teacher, while they fail to transfer the class representations. Good class representations can benefit the sample representation learning by shaping the sample representation distribution. On the other hand, the existing approaches enforce the student to fully imitate the teacher while ignoring the fact that the teacher is typically not perfect. Although the teacher has learned rich and powerful representations, it also contains unignorable bias knowledge which is usually induced by the context prior (e.g., background) in the training data. To address these two issues, in this paper, we propose comprehensive, interventional distillation (CID) that captures both sample and class representations from the teacher while removing the bias with causal intervention. Different from the existing literature that uses the softened logits of the teacher as the training targets, CID considers the softened logits as the context information of an image, which is further used to remove the biased knowledge based on causal inference. Keeping the good representations while removing the bad bias enables CID to have a better generalization ability on test data and a better transferability across different datasets against the existing state-of-the-art approaches, which is demonstrated by extensive experiments on several benchmark datasets 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7642228-d1a1-4c9b-adf3-e5a6b3f1cec4Cited by top-tier papers12
- Improved Feature Distillation via Projector EnsembleYudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu et al.NeurIPS 2022 · 73 citations
- FROSTER: Frozen CLIP is A Strong Teacher for Open-Vocabulary Action RecognitionXiaohu Huang, Hao Zhou, Kun Yao, Kai HanICLR 2024 · 56 citations
- Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisXu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao et al.NeurIPS 2022 · 49 citations
- Teach Less, Learn More: On the Undistillable Classes in Knowledge DistillationYichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu et al.NeurIPS 2022 · 42 citations
- Causal Conditional Hidden Markov Model for Multimodal Traffic PredictionYu Zhao, Pan Deng, Junting Liu, Xiaofeng Jia et al.AAAI 2023 · 17 citations
Builds on15
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
Related papers
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
- Knowledge Distillation with Refined LogitsWujie Sun, Defang Chen, Siwei Lyu, Genlang Chen et al.ICCV 2025 · 13 citations
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 1 citation
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
- Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningZijian Gao, Shanhao Han, Xingxing Zhang, Kele Xu et al.AAAI 2025 · 10 citations
