Peer Collaborative Learning for Online Knowledge Distillation
Guile Wu, Shaogang Gong
摘要
Traditional knowledge distillation uses a two-stage training strategy to transfer knowledge from a high-capacity teacher model to a compact student model, which relies heavily on the pre-trained teacher. Recent online knowledge distillation alleviates this limitation by collaborative learning, mutual learning and online ensembling, following a one-stage end-to-end training fashion. However, collaborative learning and mutual learning fail to construct an online high-capacity teacher, whilst online ensembling ignores the collaboration among branches and its logit summation impedes the further optimisation of the ensemble teacher. In this work, we propose a novel Peer Collaborative Learning method for online knowledge distillation, which integrates online ensembling and network collaboration into a unified framework. Specifically, given a target network, we construct a multi-branch network for training, in which each branch is called a peer. We perform random augmentation multiple times on the inputs to peers and assemble feature representations outputted from peers with an additional classifier as the peer ensemble teacher. This helps to transfer knowledge from a high-capacity teacher to peers, and in turn further optimises the ensemble teacher. Meanwhile, we employ the temporal mean model of each peer as the peer mean teacher to collaboratively transfer knowledge among peers, which helps each peer to learn richer knowledge and facilitates to optimise a more stable model with better generalisation. Extensive experiments on CIFAR-10, CIFAR-100 and ImageNet show that the proposed method significantly improves the generalisation of various backbone networks and outperforms the state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 被引用 103 次
- Mutual Contrastive Learning for Visual Representation LearningChuanguang Yang, Zhulin An, Linhang Cai, Yongjun XuAAAI 2022 · 被引用 95 次
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li 等CVPR 2022 · 被引用 88 次
- FedCDA: Federated Learning with Cross-rounds Divergence-aware AggregationHaozhao Wang, Haoran Xu, Yichen Li, Yuan Xu 等ICLR 2024 · 被引用 62 次
它引用的顶会 Paper4
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Mutual Mean-Teaching: Pseudo Label Refinery for Unsupervised Domain Adaptation on Person Re-identificationYixiao Ge, Dapeng Chen, Hongsheng LiICLR 2020 · 被引用 651 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu 等CVPR 2020
相关 Paper
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen 等AAAI 2023 · 被引用 16 次
- ORC: Network Group-based Knowledge Distillation using Online Role ChangeJunyong Choi, Hyeon Cho, Seokhwa Cheung, Wonjun HwangICCV 2023 · 被引用 5 次
- Weighted Mutual Learning with Diversity-Driven Model CompressionMiao Zhang, Li Wang, David Campos, Wei Huang 等NeurIPS 2022 · 被引用 10 次
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular DiversitySeonghoon Yu, Dongjun Nam, Dina Katabi, Jeany SonNeurIPS 2025 · 被引用 4 次
- Revisiting Knowledge Distillation: An Inheritance and Exploration FrameworkZhen Huang, Xu Shen, Jun Xing, Tongliang Liu 等CVPR 2021
