Adaptive Hierarchy-Branch Fusion for Online Knowledge Distillation
Linrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen, Ke Li, Ruizhi Qiao, Bo Ren, Muqing Li, Zhou Yu, Lizhuang Ma
Abstract
Online Knowledge Distillation (OKD) is designed to alleviate the dilemma that the high-capacity pre-trained teacher model is not available. However, the existing methods mostly focus on improving the ensemble prediction accuracy from multiple students (a.k.a. branches), which often overlook the homogenization problem that makes student model saturate quickly and hurts the performance. We assume that the intrinsic bottleneck of the homogenization problem comes from the identical branch architecture and coarse ensemble strategy. We propose a novel Adaptive Hierarchy-Branch Fusion framework for Online Knowledge Distillation, termed AHBF-OKD, which designs hierarchical branches and adaptive hierarchy-branch fusion module to boost the model diversity and aggregate complementary knowledge. Specifically, we first introduce hierarchical branch architectures to construct diverse peers by increasing the depth of branches monotonously on the basis of target branch. To effectively transfer knowledge from the most complex branch to the simplest target branch, we propose an adaptive hierarchy-branch fusion module to create hierarchical teacher assistants recursively, which regards the target branch as the smallest teacher assistant. During the training, the teacher assistant from the previous hierarchy is explicitly distilled by the teacher assistant and the branch from the current hierarchy. Thus, the important scores to different branches are effectively and adaptively allocated to reduce the branch homogenization. Extensive experiments demonstrate the effectiveness of AHBF-OKD on different datasets, including CIFAR-10/100 and ImageNet 2012. For example, on ImageNet 2012, the distilled ResNet-18 achieves Top-1 error of 29.28%, which significantly outperforms the state-of-the-art methods. The source code is available at https://github.com/linruigong965/AHBF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6534172a-d94f-4b21-aa88-bf87be65d5feCited by top-tier papers2
- Adaptive Dual Guidance Knowledge DistillationTong Li, Long Liu, Kang Liu, Xin Wang et al.AAAI 2025 · 1 citation
- Direction Sensitivity-Based Knowledge Distillation: Optimization-Aware Low-Rank Knowledge TransferYongkai Liao, Xinxing Chen, Zhongzheng Fu, Haoyuan Wang et al.AAAI 2026
Builds on13
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Show, Attend and Distill: Knowledge Distillation via Attention-based Feature MatchingMingi Ji, Byeongho Heo, Sungrae ParkAAAI 2021 · 194 citations
Related papers
- Peer Collaborative Learning for Online Knowledge DistillationGuile Wu, Shaogang GongAAAI 2021 · 150 citations
- Online Knowledge Distillation for Efficient Pose EstimationZheng Li, Jingwen Ye, Mingli Song, Ying Huang et al.ICCV 2021 · 123 citations
- Generalization Matters: Loss Minima Flattening via Parameter Hybridization for Efficient Online Knowledge DistillationTianli Zhang, Mengqi Xue, Jiangtao Zhang, Haofei Zhang et al.CVPR 2023
- Towards Oracle Knowledge Distillation with Neural Architecture SearchMinsoo Kang, Jonghwan Mun, Bohyung HanAAAI 2020 · 48 citations
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu et al.CVPR 2020
