Few Shot Network Compression via Cross Distillation
Haoli Bai, Jiaxiang Wu, Irwin King, Michael R. Lyu
Abstract
Model compression has been widely adopted to obtain light-weighted deep neural networks. Most prevalent methods, however, require fine-tuning with sufficient training data to ensure accuracy, which could be challenged by privacy and security issues. As a compromise between privacy and performance, in this paper we investigate few shot network compression: given few samples per class, how can we effectively compress the network with negligible performance drop? The core challenge of few shot network compression lies in high estimation errors from the original network during inference, since the compressed network can easily over-fits on the few training instances. The estimation errors could propagate and accumulate layer-wisely and finally deteriorate the network output. To address the problem, we propose cross distillation, a novel layer-wise knowledge distillation approach. By interweaving hidden layers of teacher and student network, layer-wisely accumulated estimation errors can be effectively reduced. The proposed method offers a general framework compatible with prevalent network compression techniques such as pruning. Extensive experiments n benchmark datasets demonstrate that cross distillation can significantly improve the student network's accuracy when only a few training instances are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6a69cbb5-7ce6-4d4a-9c39-db4ab7760e89Cited by top-tier papers24
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang et al.ICLR 2021 · 619 citations
- Stacked Hybrid-Attention and Group Collaborative Learning for Unbiased Scene Graph GenerationXingning Dong, Tian Gan, Xuemeng Song, Jianlong Wu et al.CVPR 2022 · 116 citations
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li et al.CVPR 2024 · 93 citations
- Towards Efficient Post-training Quantization of Pre-trained Language ModelsHaoli Bai, Lu Hou, Lifeng Shang, Xin Jiang et al.NeurIPS 2022 · 62 citations
- Compressing Transformers: Features Are Low-Rank, but Weights Are Not!Hao Yu, Jianxin WuAAAI 2023 · 59 citations
Builds on1
Related papers
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song et al.AAAI 2021 · 55 citations
- Few Sample Knowledge Distillation for Efficient Network CompressionTianhong Li, Jianguo Li, Zhuang Liu, Changshui ZhangCVPR 2020
- EPSD: Early Pruning with Self-Distillation for Efficient Model CompressionDong Chen, Ning Liu, Yichen Zhu, Zhengping Che et al.AAAI 2024 · 9 citations
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 4 citations
- Weighted Mutual Learning with Diversity-Driven Model CompressionMiao Zhang, Li Wang, David Campos, Wei Huang et al.NeurIPS 2022 · 10 citations
