Few Shot Network Compression via Cross Distillation
Haoli Bai, Jiaxiang Wu, Irwin King, Michael R. Lyu
摘要
Model compression has been widely adopted to obtain light-weighted deep neural networks. Most prevalent methods, however, require fine-tuning with sufficient training data to ensure accuracy, which could be challenged by privacy and security issues. As a compromise between privacy and performance, in this paper we investigate few shot network compression: given few samples per class, how can we effectively compress the network with negligible performance drop? The core challenge of few shot network compression lies in high estimation errors from the original network during inference, since the compressed network can easily over-fits on the few training instances. The estimation errors could propagate and accumulate layer-wisely and finally deteriorate the network output. To address the problem, we propose cross distillation, a novel layer-wise knowledge distillation approach. By interweaving hidden layers of teacher and student network, layer-wisely accumulated estimation errors can be effectively reduced. The proposed method offers a general framework compatible with prevalent network compression techniques such as pruning. Extensive experiments n benchmark datasets demonstrate that cross distillation can significantly improve the student network's accuracy when only a few training instances are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang 等ICLR 2021 · 被引用 619 次
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu 等CVPR 2022 · 被引用 116 次
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 等CVPR 2024 · 被引用 93 次
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang 等NeurIPS 2022 · 被引用 62 次
- Compressing Transformers: Features Are Low-Rank, but Weights Are Not!Hao Yu, Jianxin WuAAAI 2023 · 被引用 59 次
它引用的顶会 Paper1
相关 Paper
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song 等AAAI 2021 · 被引用 55 次
- Few Sample Knowledge Distillation for Efficient Network CompressionTianhong Li, Jianguo Li, Zhuang Liu, Changshui ZhangCVPR 2020
- EPSD: Early Pruning with Self-Distillation for Efficient Model CompressionDong Chen, Ning Liu, Yichen Zhu, Zhengping Che 等AAAI 2024 · 被引用 9 次
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 被引用 4 次
- Weighted Mutual Learning with Diversity-Driven Model CompressionMiao Zhang, Li Wang, David Campos, Wei Huang 等NeurIPS 2022 · 被引用 10 次
