Rethinking Feature-based Knowledge Distillation for Face Recognition
Jingzhi Li, Zidong Guo, Hui Li, Seungju Han, Ji-Won Baek, Min Yang, Ran Yang, Sungjoo Suh
Abstract
With the continual expansion of face datasets, featurebased distillation prevails for large-scale face recognition. In this work, we attempt to remove identity supervision in student training, to spare the GPU memory from saving massive class centers. However, this naive removal leads to inferior distillation result. We carefully inspect the performance degradation from the perspective of intrinsic dimension, and argue that the gap in intrinsic dimension, namely the intrinsic gap, is intimately connected to the infamous capacity gap problem. By constraining the teacher's search space with reverse distillation, we narrow the intrinsic gap and unleash the potential of feature-only distillation. Remarkably, the proposed reverse distillation creates universally student-friendly teacher that demonstrates outstanding student improvement. We further enhance its effectiveness by designing a student proxy to better bridge the intrinsic gap. As a result, the proposed method surpasses stateof-the-art distillation techniques with identity supervision on various face recognition benchmarks, and the improvements are consistent across different teacher-student pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0924b516-90a5-4d3e-a71f-b8e079502956Cited by top-tier papers5
- Rethinking Impersonation and Dodging Attacks on Face Recognition SystemsFengfan Zhou, Qianyu Zhou, Bangjie Yin, Hui Zheng et al.ACM MM 2024 · 9 citations
- Revisit the Essence of Distilling Knowledge through CalibrationWen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan et al.ICML 2024 · 8 citations
- Group-Wise Scaling and Orthogonal Decomposition for Domain-Invariant Feature Extraction in Face Anti-SpoofingSeungjin Jung, Kanghee Lee, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 2 citations
- TMLKD: Few-shot Trajectory Metric Learning via Knowledge DistillationDanling Lai, Jiajie Xu, Jianfeng Qu, Pingfu Chao et al.VLDB 2025 · 1 citation
- Improving the Transferability of Adversarial Attacks on Face Recognition with Diverse Parameters AugmentationFengfan Zhou, Bangjie Yin, Hefei Ling, Qianyu Zhou et al.CVPR 2025
Builds on15
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
Related papers
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 8 citations
- Evaluation-oriented Knowledge Distillation for Deep Face RecognitionYuge Huang, Jiaxiang Wu, Xingkun Xu, Shouhong DingCVPR 2022 · 35 citations
- Towards the Law of Capacity Gap in Distilling Language ModelsChen Zhang, Qiuchi Li, Dawei Song, Zheyu Ye et al.ACL 2025 · 39 citations
- Teach Less, Learn More: On the Undistillable Classes in Knowledge DistillationYichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu et al.NeurIPS 2022 · 42 citations
- Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic TeacherYong Guo, Shulian Zhang, Haolin Pan, Jing Liu et al.ICLR 2025
