Scale Decoupled Distillation
Shicai Wei, Chunbo Luo, Yang Luo
Abstract
Logit knowledge distillation attracts increasing attention due to its practicality in recent studies. However, it of-ten suffers inferior performance compared to the feature knowledge distillation. In this paper, we argue that existing log it-based methods may be sub-optimal since they only leverage the global logit output that couples multiple se-mantic knowledge. This may transfer ambiguous knowl-edge to the student and mislead its learning. To this end, we propose a simple but effective method, i.e., Scale De-coupled Distillation (SDD), for logit knowledge distillation. SDD decouples the global logit output into multi-ple local logit outputs and establishes distillation pipelines for them. This helps the student to mine and inherit fine-grained and unambiguous logit knowledge. Moreover, the decoupled knowledge can be further divided into consis-tent and complementary logit knowledge that transfers the semantic information and sample ambiguity, respectively. By increasing the weight of complementary parts, SDD can guide the student to focus more on ambiguous samples, im-proving its discrimination ability. Extensive experiments on several benchmark datasets demonstrate the effective-ness of SDD for wide teacher-student pairs, especially in the fine-grained classification task. Code is available at: https://github.comishicaiwei123/SDD-CVPR2024
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc6ea880-7c93-44f2-84e9-66bea98eb37eCited by top-tier papers3
- Cross-Architecture Distillation Made Simple with Redundancy SuppressionWeijia Zhang, Yuehao Liu, Wu Ran, Chao MaICCV 2025 · 6 citations
- VRM: Knowledge Distillation via Virtual Relation MatchingWeijia Zhang, Fei Xie, Tom Weidong Cai, Chao MaICCV 2025 · 6 citations
- Flow-Based Knowledge Transfer for Efficient Large Model DistillationXinye Yang, Junhao Wang, Rui Li, Haosen Sun et al.AAAI 2026
Builds on6
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang et al.AAAI 2021 · 368 citations
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou et al.ICLR 2021 · 209 citations
- From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsZhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang et al.ICCV 2023 · 141 citations
Related papers
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- Multi-Level Logit DistillationYing Jin, Jiaqi Wang, Dahua LinCVPR 2023
- Distilling Global and Local Logits with Densely Connected RelationsYoumin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali et al.ICCV 2021 · 33 citations
- SDE : Scale-Difference Evolution Knowledge DistillationHejie LuKDD 2026
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng et al.ICCV 2023 · 16 citations
