Compressing Models with Few Samples: Mimicking then Replacing
Huanyu Wang, Junjie Liu, Xin Ma, Yang Yong, Zhenhua Chai, Jianxin Wu
Abstract
Few-sample compression aims to compress a big redundant model into a small compact one with only few samples. If we fine-tune models with these limited few samples directly, models will be vulnerable to overfit and learn almost nothing. Hence, previous methods optimize the compressed model layer-by-layer and try to make every layer have the same outputs as the corresponding layer in the teacher model, which is cumbersome. In this paper, we propose a new framework named Mimicking then Replacing (MiR) for few-sample compression, which firstly urges the pruned model to output the same features as the teacher's in the penultimate layer, and then replaces teacher's layers before penultimate with a well-tuned compact one. Unlike previous layer-wise reconstruction methods, our MiR optimizes the entire network holistically, which is not only simple and effective, but also unsupervised and general. MiR outperforms previous methods with large margins. Codes is available at https://github.com/cjnjuwhy/MiR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7cddb8ad-dabe-4cf5-a3b4-ab7168a0aff6Cited by top-tier papers4
- Compressing Transformers: Features Are Low-Rank, but Weights Are Not!Hao Yu, Jianxin WuAAAI 2023 · 59 citations
- Dense Vision Transformer Compression with Few SamplesHanxiao Zhang, Yifan Zhou, Guo-Hua WangCVPR 2024 · 5 citations
- Practical Network Acceleration with Tiny SetsGuo-Hua Wang, Jianxin WuCVPR 2023
- SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuningMinjun Kim, Jongjin Kim, U KangICLR 2025
Builds on11
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- Data-Free Learning of Student NetworksHanting Chen, Yunhe Wang, Chang Xu, Zhaohui Yang et al.ICCV 2019 · 427 citations
- Group Fisher Pruning for Practical Network CompressionLiyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou et al.ICML 2021 · 204 citations
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 66 citations
- Progressive Network Grafting for Few-Shot Knowledge DistillationChengchao Shen, Xinchao Wang, Youtan Yin, Jie Song et al.AAAI 2021 · 55 citations
Related papers
- CompRess: Self-Supervised Learning by Compressing RepresentationsSoroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed PirsiavashNeurIPS 2020 · 105 citations
- Few Sample Knowledge Distillation for Efficient Network CompressionTianhong Li, Jianguo Li, Zhuang Liu, Changshui ZhangCVPR 2020
- Knowledge distillation via softmax regression representation learningJing Yang, Brais Martínez, Adrian Bulat, Georgios TzimiropoulosICLR 2021 · 55 citations
- Go Wide, Then Narrow: Efficient Training of Deep Thin NetworksDenny Zhou, Mao Ye, Chen Chen, Tianjian Meng et al.ICML 2020 · 21 citations
- Beyond Student: An Asymmetric Network for Neural Network InheritanceYiyun Zhou, Jingwei Shi, Mingjing Xu, Zhonghua Jiang et al.ICLR 2026 · 1 citation
