Weighted Mutual Learning with Diversity-Driven Model Compression
Miao Zhang, Li Wang, David Campos, Wei Huang, Chenjuan Guo, Bin Yang
摘要
Online distillation attracts attention from the community as it simplifies the traditional two-stage knowledge distillation process into a single stage. Online distillation collaboratively trains a group of peer models, which are treated as students, and all students gain extra knowledge from each other. However, memory consumption and diversity among students are two key challenges to the scalability and quality of online distillation. To address the two challenges, this paper presents a framework called Weighted Mutual Learning with Diversity-Driven Model Compression ( WML ) for online distillation. First, at the base of a hierarchical structure where students share different parts, we leverage the structured network pruning to generate diversified students with different models sizes, thus also helping reduce the memory requirements. Second, rather than taking the average of students, this paper, for the first time, leverages a bi-level formulation to estimate the relative importance of students with a close-form, to further boost the effectiveness of the distillation from each other. Extensive experiments show the generalization of the proposed framework, which outperforms existing online distillation methods on a variety of deep neural networks. More interesting, as a byproduct, WML produces a series of students with different model sizes in a single run, which also achieves competitive results compared with existing channel pruning methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper24
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
相关 Paper
- Peer Collaborative Learning for Online Knowledge DistillationGuile Wu, Shaogang GongAAAI 2021 · 被引用 150 次
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Online Knowledge Distillation via Collaborative LearningQiushan Guo, Xinjiang Wang, Yichao Wu, Zhipeng Yu 等CVPR 2020
- Adaptive Hierarchy-Branch Fusion for Online Knowledge DistillationLinrui Gong, Shaohui Lin, Baochang Zhang, Yunhang Shen 等AAAI 2023 · 被引用 16 次
- Online Knowledge Distillation for Efficient Pose EstimationZheng Li, Jingwen Ye, Mingli Song, Ying Huang 等ICCV 2021 · 被引用 123 次
