Grouped Knowledge Distillation for Deep Face Recognition
Weisong Zhao, Xiangyu Zhu, Kaiwen Guo, Xiaoyu Zhang, Zhen Lei
摘要
Compared with the feature-based distillation methods, logits distillation can liberalize the requirements of consistent feature dimension between teacher and student networks, while the performance is deemed inferior in face recognition. One major challenge is that the light-weight student network has difficulty fitting the target logits due to its low model capacity, which is attributed to the significant number of identities in face recognition. Therefore, we seek to probe the target logits to extract the primary knowledge related to face identity, and discard the others, to make the distillation more achievable for the student network. Specifically, there is a tail group with near-zero values in the prediction, containing minor knowledge for distillation. To provide a clear perspective of its impact, we first partition the logits into two groups, i.e., Primary Group and Secondary Group, according to the cumulative probability of the softened prediction. Then, we reorganize the Knowledge Distillation (KD) loss of grouped logits into three parts, i.e., Primary-KD, Secondary-KD, and Binary-KD. Primary-KD refers to distilling the primary knowledge from the teacher, Secondary-KD aims to refine minor knowledge but increases the difficulty of distillation, and Binary-KD ensures the consistency of knowledge distribution between teacher and student. We experimentally found that (1) Primary-KD and Binary-KD are indispensable for KD, and (2) Secondary-KD is the culprit restricting KD at the bottleneck. Therefore, we propose a Grouped Knowledge Distillation (GKD) that retains the Primary-KD and Binary-KD but omits Secondary-KD in the ultimate KD loss calculation. Extensive experimental results on popular face recognition benchmarks demonstrate the superiority of proposed GKD over state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Cross-Architecture Distillation for Face RecognitionWeisong Zhao, Xiangyu Zhu, Zhixiang He, Xiaoyu Zhang 等ACM MM 2023 · 被引用 10 次
- KeyPoint Relative Position Encoding for Face RecognitionMinchul Kim, Yiyang Su, Feng Liu, Anil Jain 等CVPR 2024
它引用的顶会 Paper7
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- Nested Collaborative Learning for Long-Tailed Visual RecognitionJun Li, Zichang Tan, Jun Wan, Zhen Lei 等CVPR 2022 · 被引用 95 次
相关 Paper
- Multi-Level Logit DistillationYing Jin, Jiaqi Wang, Dahua LinCVPR 2023
- Scale Decoupled DistillationShicai Wei, Chunbo Luo, Yang LuoCVPR 2024 · 被引用 32 次
- Streamlined Knowledge DistillationHyeon-Jin Jeong, Han-Jin Lee, Seok-Hwan ChoiCVPR 2026 · 被引用 1 次
- Distilling Balanced Knowledge from a Biased TeacherSeonghak KimCVPR 2026 · 被引用 1 次
- ICD-Face: Intra-class Compactness Distillation for Face RecognitionZhipeng Yu, Jiaheng Liu, Haoyu Qin, Yichao Wu 等ICCV 2023 · 被引用 7 次
